Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty. Customers order around the clock, suppliers reprice and fail, and market shocks arrive with partial or no warning. The agent acts through the same 29 merchant tools a human operator would use, under a windowed operation budget that makes simulated time a function of actions taken, so model latency cannot influence simulated time. Pass thresholds are calibrated against scripted anchor policies, the reward is hardened against a catalogue of reward hacks, and every episode replays identically given a sequence of actions. We evaluate seven frontier LLMs on 11 scenarios of 30 to 45 days and a full simulated year, over three world seeds at matched reasoning effort. No model matches the scripted smart-triage policy on average: the best, DeepSeek-V4-Pro, passes 49% of task-seed cells against the heuristic's 97%. Human experts working through the same tools and budgets outscore every model (mean composite 0.708 vs. 0.700). Over a full simulated year under the Claude Code harness, most models show dramatic performance improvement. In a GRPO post-training run, Qwen3.5-27B trained on only five disjoint tasks raises its mean composite on the held-out evaluation tasks from 0.136 to 0.373. We release five example training-split tasks, ten sample trajectories, and the scoring and verification tooling; the full environment and evaluation suite are withheld to keep the benchmark uncontaminated.
Figures & tables
Figure 1: StoreBench overview. A harness-agnostic commerce agent reaches the world only through 29 MCP tools; each call spends one operation, and the world advances in 12-hour windows. The Ostrelle store and its supplier market run on live Medusa v2 software, while demand, order flow, scripted events, and rivals sit in a hidden market the agent observes only through its own sales. Episodes of 30 to 365 days receive weighted scores and are compared to scripted anchor policies. Human experts operate the same MCP surface through a thin GUI, under the same budgets and grading. Icons: Microsoft Fluent Emoji (MIT License).
Figure 2: Three example tasks from the 12-task suite: steady (steady-state operations), trade-war (a correlated supply shock), and black-swan (a silent demand crash). Each chart plots the hidden quantity the task’s event script perturbs; markers show when each event fires and through which channel the agent can learn of it. Quotes are from the task prompts.
Policy / Model
Composite
Pass
Business
Reliab.
Contin.
Smart-triage
0.764
97%
–
–
–
Rule-based
0.638
12%
–
–
–
Blast-list
0.463
0%
–
–
–
Absentee
0.100
0%
–
–
–
Do-nothing
0.000
0%
–
–
–
Exploiter
0.039
0%
–
–
–
Table 1: StoreBench leaderboard: we report mean composite score, pass rate, and component means (business, reliability, continuity per Eq. 1 ) over the 11 eval tasks × 3 seeds × 3 attempts, at matched reasoning effort; a trial passes if its composite clears its cell’s calibrated threshold (§ 4.2 ). Deterministic policies are italicized; no model reaches the smart-triage operator on average. Bold marks the best model or human value per column. Dashes: component means are not reported for scripted policies.
Model
Short
Full-yr
Δ
DeepSeek-V4-Pro
0.700
0.976
+0.276
Qwen3.8-Max
0.587
0.974
+0.387
Claude Opus 4.8
0.668
0.971
+0.303
Claude Fable 5
0.629
0.966
+0.337
Muse Spark 1.2
0.387
0.862
+0.475
GPT-5.6 Sol
0.538
0.783
+0.245
Table 2: Long horizon vs. short horizon. Δ is full-year minus the eval mean; the setups differ in harness.
Policy
Composite
Pass
Business
Reliab.
Contin.
On-time
Returns
Windows
Base
0.136
0/99
0.005
0.401
0.521
0.65
0.16
33.4
Tuned
0.373
3/99
0.236
0.669
0.741
0.89
0.45
47.3
Table 3: Post-training transfers to held-out tasks. Qwen3.5-27B before and after GRPO on the 11 evaluation tasks × 3 seeds × 3 episodes, official composite score, and its components. On-time fulfillment rates and successful return rates are also shown. Both policies run under the same harness.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Scenario
Days
Windows
Ops/win
Events
Competition
steady
Steady-state operations
30
60
20
1
–
cold-start
Rebuild the shop
30
60
12
1
–
supply-shock
Your biggest supplier fails
30
60
20
4
–
quality-crisis
Find the bad batch
30
60
20
2
–
liquidity
Run it on fumes
30
60
11
0
–
peak-season
The platform festival
30
60
16
2
–
Appendix
Table 4: The StoreBench task suite. All tasks share the seeded world; each lens sets horizon, operation budget, and event script.
Model
0.7/0.2/0.1
0.5/0.4/0.1
Business only
DeepSeek-V4-Pro
0.700
0.748
0.614
Claude Opus 4.8
0.668
0.720
0.573
Claude Fable 5
0.629
0.689
0.520
Qwen3.8-Max
0.587
0.655
0.465
Gemini 3.8 Flash
0.552
0.608
0.440
GPT-5.6 Sol
0.538
0.619
0.398
Appendix
Table 5: Mean composite under alternative weightings (business/reliability/continuity). The ranking is essentially stable across weightings.
Task
BL
RB
ST
Threshold
steady
0.49/0.51/0.45
0.67/0.63/0.66
0.77/0.77/0.77
0.72/0.70/0.72
cold-start
0.10/0.10/0.10
0.13/0.13/0.13
0.74/0.75/0.74
0.43/0.44/0.44
supply-shock
0.49/0.53/0.45
0.69/0.66/0.63
0.77/0.77/0.77
0.73/0.71/0.70
quality-crisis
0.50/0.52/0.43
0.71/0.65/0.65
0.77/0.77/0.77
0.74/0.71/0.71
liquidity
0.50/0.58/0.52
0.75/0.72/0.76
0.75/0.76/0.75
0.74/0.74/0.74
peak-season
0.45/0.50/0.52
0.69/0.69/0.65
0.76/0.76/0.77
0.73/0.73/0.71
Appendix
Table 6: Anchor composites and calibrated pass thresholds for every task-seed cell (seed 7 / 11 / 13). BL = blast-list, RB = rule-based, ST = smart-triage; the threshold separates the top two honest anchors under the margin and tie gates of § 4.2 .
Figure 3: The anchor ladder, measured per task-seed cell (one marker per seed). The calibrated pass threshold sits between the two best honest anchors on every cell; cold-start shows the ladder at its widest, with checklist play barely above volume play and the threshold set by the tie-gate against the rung below.
Task
DeepS
Opus
Fable
Qwen
Gemini
GPT
Muse
Human
ST
steady
0.736
0.583
0.603
0.585
0.533
0.511
0.290
0.539
0.770
cold-start
0.851
0.632
0.855
0.767
0.545
0.466
0.692
0.794
0.741
supply-shock
0.638
0.673
0.493
0.459
0.449
0.555
0.295
0.626
0.770
quality-crisis
0.696
0.665
0.680
0.594
0.440
0.432
0.256
0.614
0.772
liquidity
0.723
0.689
0.670
0.754
0.513
0.645
0.552
0.733
0.753
peak-season
0.748
0.710
0.783
0.649
0.659
0.689
0.501
0.760
0.765
Appendix
Table 7: Mean composite scores by model and task (3 seeds × 3 attempts per cell). Human is the human expert’s score; ST is the smart-triage heuristic anchor for reference. full-year is measured under a separate agentic harness with extended thinking and is reported separately below eval scores.
Figure 4: Short-horizon mean (open marker, 11 tasks) versus the 365-window full-year task (filled marker). Most models operate well above smart-triage over the full year; the two with observed long-horizon failure modes (Gemini, GPT) are in orange. Dotted line: smart-triage.
Task
Base
Step 34
Step 36
Δ36
BL
RB
steady
0.129
0.353
0.574
+0.445
0.481
0.656
cold-start
0.110
0.089
0.082
−0.028
0.100
0.133
supply-shock
0.134
0.289
0.353
+0.220
0.488
0.659
quality-crisis
0.127
0.435
0.382
+0.255
0.483
0.669
liquidity
0.205
0.573
0.653
+0.449
0.536
0.741
peak-season
0.155
0.443
0.409
+0.254
0.492
0.674
Appendix
Table 8: Post-training by held-out task: mean official composite over 3 seeds × 3 episodes for the base model and GRPO checkpoints 34 and 36, with the blast-list (BL) and rule-based (RB) anchors on the same cells.