ShelfLife E-Sim

A stateful simulator of retail operations.

Tasks
12
Last Measured
Aug 26, 2026

E-Sim is a long-horizon resource management benchmark built on a stateful simulator of retail operations. Agents take control of the business on a fixed date and operate it one day at a time. Environments are grounded in nearly a decade of real customer orders, purchase orders, and inventory snapshots from a live retailer.

Overall Leaderboard

Model performance ranked across tasks

Model
Mean Reward (pts)
Cost / Rollout
Output Tokens
Grok 4.6
Likely mean reward (pts)63%65%67%
65.0
± 1.7 pts
$5.88
57.9k
GPT-5.6 Sol
Likely mean reward (pts)52%54%56%
54.1
± 2.2 pts
$5.06
39.8k
Gemini 3.7 Flash
Likely mean reward (pts)50%52%54%
51.7
± 2.2 pts
$2.44
55.3k
Fable 5
Likely mean reward (pts)49%51%53%
51.0
± 2.3 pts
$8.45
54.3k
Opus 5
Likely mean reward (pts)37%40%42%
39.5
± 2.1 pts
$4.81
55.6k
Kimi K3
Likely mean reward (pts)35%38%40%
37.6
± 2.8 pts
$4.42
50.6k

Skill Profile

Each model's performance on the tasks tagged with a given skill.

25%50%75%100%EconomicExecutionInventoryTurnoverCommitmentControlInventory-PolicyRobustnessSelf-EvaluationCalibration
Grok 4.6GPT-5.6 SolGemini 3.7 FlashFable 5Opus 5Kimi K3

Skill Leaderboard

The same figures as a table, with the best score in each row marked.

SkillGrok 4.6GPT‑5.6 SolGemini 3.7 FlashFable 5Opus 5Kimi K3
Economic executionThe model's overall ability to execute a long-horizon resource management task.92.7%79.2%78.1%75.0%57.9%47.9%
Inventory turnoverWhether the model converts the stock it carries into sales rather than paying to hold it: share of rollouts at or below the task's median holding-cost-to-revenue ratio.78.4%43.8%64.6%64.6%21.9%27.0%
Commitment controlModel's ability to avoid ending with excessive unpaid purchase-order commitments.28.6%70.0%52.4%33.3%32.5%28.9%
Inventory-policy robustnessWhether the model buys and moves the right inventory to the right locations at the right time.97.8%84.4%87.2%90.0%80.9%76.1%
Self-evaluation calibrationModel's ability to avoid optimistic claims that ignore benchmark performance.71.4%90.0%52.4%33.3%82.5%52.6%

Cost / Performance

Score against dollars per rollout. The dotted line is the Pareto frontier.

1020304050607080$0$2$4$6$8$10Pareto frontierGrok 4.6GPT-5.6 SolGemini 3.7 FlashFable 5Opus 5Kimi K3More efficientCost / RolloutMean Reward (pts)