ShelfLife

A digital twin of a real e-commerce company.

Tasks
200
Last Measured
Aug 13, 2026

We created a digital twin, ShelfLife, of a live e-commerce company. From ShelfLife, we present a collection of 200 environments drawn from the company's operational data over the past decade to answer the question: can you deploy an agent in a leadership role within this company, and trust it to execute with appropriate judgment and precision?

Every environment is constructed in direct collaboration with the company's leaders and the operators who performed the work in production. We worked iteratively with them to develop a deep understanding of their workflows and domain expertise.

We report benchmark results for the following models on this collection: Claude Opus 5, Claude Fable 5, GPT-5.6 Sol, Grok 4.6, Kimi K3, and Gemini 3.6 Flash.

Overall Leaderboard

Share of tasks each model passed on a single attempt.

Model
Pass@1
Cost / Rollout
Opus 5
Likely pass@168%69%70%
69.0%
± 0.6 pp
$5.02
GPT-5.6 Sol
Likely pass@154%56%58%
55.7%
± 1.9 pp
$2.30
Fable 5
Likely pass@154%55%55%
54.7%
± 0.8 pp
$17.04
Grok 4.6
Likely pass@139%39%40%
39.4%
± 0.7 pp
$3.61
Kimi K3
Likely pass@130%31%31%
30.6%
± 0.9 pp
$7.83
Gemini 3.6 Flash
Likely pass@116%16%17%
16.1%
± 0.5 pp
$2.29

Skill Profile

Each model's pass rate on the tasks tagged with a given skill.

25%50%75%100%Long-HorizonPlanningDerivingUnstated MethodsNumericalFinancePrecisionOptimizationUnder ConstraintLong-HorizonState ManagementSafety AndDisciplineIdentifyingQuality Of DataComplex Tool-UseAvoidingHallucination InArtifactsDisclosureControl
Opus 5GPT-5.6 SolFable 5Grok 4.6Kimi K3Gemini 3.6 Flash

Skill Leaderboard

The same figures as a table, with the best score in each row marked.

SkillTasksOpus 5GPT‑5.6 SolFable 5Grok 4.6Kimi K3Gemini 3.6 Flash
Long-horizon planningTasks in this collection tagged “Long-horizon planning” — 87 of them.8778.8%60.8%65.1%28.7%22.6%1.7%
Deriving unstated methodsTasks in this collection tagged “Deriving unstated methods” — 106 of them.10640.3%38.0%43.7%21.9%15.6%7.1%
Numerical finance precisionTasks in this collection tagged “Numerical finance precision” — 141 of them.14167.2%50.5%47.7%43.5%30.0%13.4%
Optimization under constraintTasks in this collection tagged “Optimization under constraint” — 46 of them.4663.1%24.9%34.3%35.3%19.6%12.2%
Long-horizon state managementTasks in this collection tagged “Long-horizon state management” — 31 of them.3168.5%69.4%62.5%29.3%48.3%60.2%
Safety and disciplineTasks in this collection tagged “Safety and discipline” — 59 of them.5985.8%55.8%60.4%9.9%25.1%14.7%
Identifying quality of dataTasks in this collection tagged “Identifying quality of data” — 33 of them.3354.3%39.2%77.7%17.8%35.0%26.3%
Complex tool-useTasks in this collection tagged “Complex tool-use” — 13 of them.1398.1%88.5%100.0%94.9%75.0%0.0%
Avoiding hallucination in artifactsTasks in this collection tagged “Avoiding hallucination in artifacts” — 7 of them.726.8%25.8%26.8%13.3%25.4%13.1%
Disclosure controlTasks in this collection tagged “Disclosure control” — 27 of them.2753.8%54.5%49.5%25.2%47.5%35.5%
Overall20069.0%55.7%54.7%39.4%30.6%16.1%

Cost / Performance

Score against dollars per rollout. The dotted line is the Pareto frontier.

10%20%30%40%50%60%70%80%$0$5$10$15$20Pareto frontierOpus 5GPT-5.6 SolFable 5Grok 4.6Kimi K3Gemini 3.6 FlashMore efficientCost / RolloutPass@1