ShelfLife E-Sim

A stateful simulator of retail operations.

Tasks
12
Last Measured
Aug 26, 2026

E-Sim is a long-horizon resource management benchmark built on a stateful simulator of retail operations. Agents take control of the business on a fixed date and operate it one day at a time. Environments are grounded in nearly a decade of real customer orders, purchase orders, and inventory snapshots from a live retailer.

ShelfLife E-Sim: Documentation

E-Sim is a long-horizon resource management benchmark built on a stateful simulator of retail operations. An agent takes control of the business on a fixed date and operates the business one day at a time with tools.

While benchmarks like Vending Bench grade decisions with synthetic data, E-Sim is grounded in nearly a decade of real customer orders, purchase orders, and inventory snapshots from a live retailer. Decisions made in E-Sim are compared to deterministic policies using the same historical demand.

Unlike traditional finance or retail benchmarks like GDPval, E-Sim tasks evaluate inventory control and capital commitment in a stateful environment with hidden information, random events, delayed effects, and irreversible actions under a fixed turn budget.

Task Overview

In E-Sim tasks, agents must manage purchasing and inventory allocation for a retailer over a number of days. Demand follows a hidden record of historical orders. Each day, customer orders will arrive and the agent must decide three things: what to restock, when to buy it, and where to hold the stock. Agents must also balance sales, purchasing capacity, transfer fees, returns, and holding cost.

Each run starts with conditions based on real data:

  • Inventory from a historical snapshot of the company, divided between a main warehouse and eight stores.
  • Starting cash from a median historical Purchase Order (PO) spend.
  • Real supplier POs and unpaid balances.

Simulation and Tools

Agents being evaluated with E-Sim tasks cannot see future demand, but can inspect historical past sales, products, supplier prices, purchase orders, inventory reports, and operational notes shipped with each task. The harness gives each model access to bash and read tools for multiple file formats.

Time in the simulation advances via a sleep tool (called with an end date parameter), which will tick the simulation forward until the end date. The tool can also be configured by the agent to wake for specified events, such as a shipment arrival or stockout.

Any other tool call does not advance the simulation tick. The agent submits purchase orders and transfers in the morning with their respective tools. The simulation puts them in a queue and reserves the necessary cash or stock immediately. The orders and the transfers become effective when the sleep tool is called and ends the day.

The simulation processes each closed day in this sequence:

  1. Supplier shipments that are due on that day arrive at the warehouse.
  2. The simulation pays the supplier balances that are due on that day.
  3. Store transfers that are due on that day arrive.
  4. The simulation dispatches the POs and the transfers from the morning queue.
  5. Each location supplies the customer demand of that day for its own orders.
  6. The simulation refunds and restocks the returns from the sales of ten days before.
  7. The simulation charges holding cost on the inventory at the end of the day.

E-Sim also contains a transfer tool for moving inventory between any two locations. A transfer takes at least one day, costs money for each unit it sends, and has a probability to arrive short (i.e. missing a percentage of units). Each sold unit also has some chance of being returned by the customer ten days after the sale.

All rollouts were run on the model provider’s preferred harness with high effort, e.g. Grok Build CLI for Grok 4.6, Claude Code for Anthropic models, Codex for GPT-5.6 Sol, and Antigravity CLI for Gemini 3.7 Flash.

Randomness

E-Sim tasks have three types of random outcomes

  1. Supplier delays. Each PO from one specific supplier can arrive anywhere from 1 to 10 days after its quoted arrival date.
  2. Transfer shortages. Each inventory transfer has a chance of arriving short.
  3. Customer returns. Each sold unit has a chance of being returned.

Historical demand and estimated missing demand have no random/nondeterministic component. E-Sim uses seeded pseudo-randomness. For each possible outcome, the simulation calculates a random number from a task-wide base seed, event ID, and draw number. For example, a transfer shortage uses a base seed, the name transfer-shortage, and the transfer ID. A customer return uses a separate draw number for each unit.

Reconstructing Demand

E-Sim simulations evaluate the demand one order line at a time by replaying historical order traces of real customer orders. For each order line that is filled, the simulation credits revenue to the agent. If a location doesn’t have the full quantity for that order line, the agent loses the full order. We choose not to reveal the lost orders to the agent.

However, simply replaying historical orders as a proxy for customer demand is inaccurate, because the demand signal is corrupted by stockouts. A day with zero orders for a product could mean either nobody wanted to order it, or someone couldn’t buy it because the store didn’t have it in stock. A naive implementation of the simulation would unfairly penalize an agent that restocks a Stock Keeping Unit (SKU) the real shop let become empty (the corrected stockout gives no sales and the purchase would register as an over-order).

To avoid this, E-Sim conservatively estimates demand from verified inventory records. We first identify days when an item had no sales. We use inventory snapshots, shipment records, and sales records to confirm that the SKU was out of stock under every plausible arrival schedule.

For each stockout day, we estimate missing sales from the SKU’s in-stock history, move weak estimates toward the average sales rate across all SKUs, weight it by the month and weekday, then assign the units to locations from historical sales. The sales rate is adjusted with gamma-Poisson shrinkage based on seasonally-weighted in-stock days. SKUs with little historical recorded sales depend mostly on the average sales rate. If the SKU has many historical recorded sales, the estimated rate depends mostly on its own sales rate.

Across the full dataset used to build an E-Sim task’s demand, real historical customer orders account for about 99.6% of units. The remaining 0.4% of units sold is estimated demand.

Evaluation

E-Sim tasks ship with a hidden grader that examines a run from a saved state file. The grader calculates an adjusted cash score. This score is the change in cash after the grader subtracts unpaid PO balances and lost-sale penalties. Lost sale penalties are deductions to the cash score that are equal to a percentage of each order line the agent was unable to fill. For each unfilled line, the penalty deduction is the configured penalty rate multiplied by the quantity and unit price.

A submission scores more optimally if the run is complete and if the agent’s adjusted cash score is strictly greater than that of a reference buyer.

The reference buyer is a deterministic replay against the simulation that sets the reward baseline. Its strategy is a greedy purchasing algorithm that uses a forecast based on the previous 90 days. It selects the least expensive supplier that can deliver on time. It does not know future demand or random outcomes.

Reward

The grader calculates a reward from two deterministic replays of the same seed. Rewards are scaled against a reference buyer, and confirmed to be fair.

The reward scale uses 2 deterministic benchmark scores as the floor and ceiling:

  1. A floor. The buyer does nothing. The reward is 0.10.
  2. An aggregate demand buyer. This buyer knows the total demand for each SKU at each location during the simulation. It does not know the dates of specific orders or the results of random events. It adds a safety margin that balances the cost of excess inventory against the cost of a stockout (a newsvendor quantile). The buyer distributes the resulting demand target across the simulation days according to the share of total demand that occurs on each day across all SKUs. This score represents a ceiling that a policy with accurate demand-rate forecasts could approach. The reward is 0.90.

Between the floor and the distributional value, the reward increases at a constant rate. Each additional unit of adjusted cash score has the same reward value in this range.

reward = 0.10 + 0.80 × (adjustedCash - floor) / (distributional - floor)

Scores above the distributional value increase toward 1.0 at a slower rate.

Results

Across graded rollouts, Grok 4.6 achieves the highest mean reward (65.0). GPT-5.6 Sol achieves the highest single-task mean reward (0.65). The bottom of the leaderboard is Opus 5 (39.5) and Kimi K3 (37.6). Gemini 3.7 Flash was the most cost-efficient model, averaging $2.44 per rollout at standard (non-introductory) API rates.

We analyzed rollout transcripts to understand model failures with an automated judge, which identified the following behaviors:

Gemini is the most paranoid, Opus 5 is the least

All models use the sleep tool as an event-aware way to control timing in the simulation. The models mainly differ on whether they request long or short sleep horizons. Kimi K3 had the longest median planned horizon of 5 days, while Gemini 3.7 Flash had the shortest, at 1 day.

Gemini 3.7 Flash advances the simulation nearly one day at a time (123 sleeps per single-task run, median planned horizon of 1 day) and subscribes to nearly every wake event on nearly every call.

Opus 5 is the least paranoid, with the fewest sleeps (29-44 per run), long sleep horizons (median 4 days), and by far the narrowest wake event conditions (1.4 of the 6 available, versus 1.9-4.6 for every other model). Uniquely, Opus never subscribes to the payable_paid event anywhere in the analyzed rollouts.

The best-performing model, Grok 4.6, barely monitors interactively. Instead, it batches multi-day advances in scripted bash loops. Vigilance/paranoia is a weak performance predictor in E-Sim: within a task, how many wake conditions a model subscribes to per sleep is uncorrelated with task performance. Instead, the winning postures are more selective, such as Grok’s minimal scripted sleep orchestration, which outperforms Gemini (the most vigilant model).

Unpaid commitments are the dominant failure mechanism

The largest failure mechanism in the corpus is terminal commitment: 60% of all completed failures end with more than $100k of unpaid purchase-order balances, and in 44% of failures the verdict flips to a pass if the commitment is added back. With Opus 5, unpaid purchase orders at the end of the simulation represent 68% of its runs that scored below the reference buyer. The grader subtracts unpaid commitments from the final net cash value, so high ending cash does not always mean the rollout performed well. If the commitment was added back (i.e. the model did not leave unpaid commitments), the rollout’s adjusted cash would exceed the reference buyer’s in 48% of cases.

Some models treated incoming inventory as a benefit, even when it would arrive after the scoring date. An open purchase order uses future cash, and a manager cannot treat that cash as available while ignoring the amount still owed to the supplier. Inventory that arrives after the review date also cannot prevent earlier lost sales.

In Opus 5 specifically, 68% of runs that underperform the reference buyer are due to material commitments. Same for Kimi K3, with 75% of its failures due to material commitments, the highest share of any model. Its transcripts show the same two mechanics as Opus: net-terms balances are treated as settled and payables are deliberately deferred past the scoring date. Beyond these two, Fable 5’s commitment discipline degrades at scale: its median commitment ($205k) is the worst in all benchmarked models, including the single largest observed commitment ($1.56M). Only GPT-5.6 Sol is avoids this failure mode, with a median commitment of $57k, and no commitment over $500k in any of its rollouts.

Models restock the wrong SKU, location, quantity, or arrival timing

In 38% of completed failures, the model completes the horizon, actively buys and transfers inventory, ends with modest commitments and transfer costs, and still loses to the reference buyer. In these runs the inventory policy selected the wrong combination of SKU, location, quantity, or arrival timing. This is partially due to the design of the task. E-Sim evaluates each order line against inventory for the order’s specific location and SKU. If that location lacks the full quantity, the whole order line counts as a lost sale. Across these runs, all models incurred at least $1M in lost sale penalties, whereas the median reference policy shortfall was ~$266k.