Long-Horizon Scaling: How Model Capabilities Shape the Returns to Computation
Organizations: Meituan LongCat Team · Zhejiang University · Westlake University · Allen Institute for AI · University of Washington
Abstract
Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab and EdgeBench, two long-horizon benchmarks. We find that starting performance and subsequent growth are associated with different capabilities: within a task category, similar early scores can precede different later gains. To formalize this finding, we model capability-time scaling with category-specific logistic power laws shared across models. Fitted to early trajectories, these curves extrapolate the observed models' category-average scores to later computation. However, rising average scores mask narrowing improvement opportunities: later gains concentrate among fewer improving models. High final scores and continued improvement also have distinct capability profiles. Predicted mean gains estimate each model's fraction of improving tasks; averaging these estimates forecasts the average share of improving models. These uneven returns motivate deciding whether a specific run should continue. We therefore derive a continuation policy to save time and compute with limited score loss. The policy conditions growth predictions on the run's observed progress and weighs immediate and delayed gains against computation costs. In replay with training and price calibration based on other models' histories, the policy saves roughly one-third of full-run time, with relative score losses of 2.4% on AutoLab individual runs and 3.3% on EdgeBench published mean curves. Our repository is available at https://github.com/Chihaya-Anon-chan/long-horizon-scaling.
Figures & tables
| AutoLab | EdgeBench | ||||
|---|---|---|---|---|---|
| Matched condition | Reported outcome | Fixed budget | Ours | Fixed budget | Ours |
| Save time | Score retained (%) | ||||
| Save time | Score retained (%) | ||||
| Retain score | Time saved (%) | ||||
| Retain score | Time saved (%) | ||||
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Number of PCs | 1 | 3 | 5 | 7 | 29 |
|---|---|---|---|---|---|
| Mean later-score RMSE | 4.861 | 3.180 | 2.315 | 2.288 | 2.504 |
| Category | Affine in log time | Logistic power law |
|---|---|---|
| CUDA | 4.42 | 4.02 |
| Model Development | 5.68 | 5.73 |
| Puzzle & Challenge | 5.86 | 4.28 |
| System Optimization | 3.91 | 3.48 |
| Formal | 2.03 | 1.53 |
| Games | 1.09 | 1.07 |
| Benchmark | Early coefficient fit | Later evaluation |
|---|---|---|
| AutoLab | ||
| EdgeBench | h | h |
| Category | Last score | Independent | Indicators | PCs |
|---|---|---|---|---|
| CUDA | 4.19 | 4.95 | 4.18 | 3.75 |
| Model Development | 6.19 | 3.46 | 6.06 | 3.79 |
| Puzzle & Challenge | 5.83 | 4.27 | 4.41 | 4.08 |
| System Optimization | 1.71 | 3.97 | 3.41 | 3.46 |
| Formal | 4.40 | 1.86 | 1.66 | 1.93 |
| Games | 1.95 | 1.10 | 1.01 | 1.03 |
| Fixed primary panel | Expanded panel | |||
| Category | Tasks | Decline (pp) | Tasks | Decline (pp) |
| CUDA | 3 | 4 | ||
| Model Development | 2 | 2 | ||
| Puzzle & Challenge | 5 | 9 | ||
| System Optimization | 12 | 14 | ||
| Formal | 8 | 8 | ||
| Category | Models/tasks | Task process |
|---|---|---|
| CUDA | 4/3 | Optimize decoding, geometric correspondence, and elliptic-curve CUDA kernels. |
| Model Development | 4/2 | Select fine-tuning data and optimize an online serving engine. |
| Puzzle & Challenge | 4/5 | Reduce program or model complexity under correctness and accuracy constraints. |
| System Optimization | 4/12 | Improve implementations of cryptography, retrieval, storage, and numerical kernels. |
| Formal | 5/8 | Complete interdependent Lean/Coq proofs in multi-file projects. |
| Games | 5/8 | Explore interactive fiction or implement game-playing and management agents. |
| Category | Final score | Terminal frequency | Difference |
|---|---|---|---|
| CUDA | |||
| Model Development | |||
| Puzzle & Challenge | |||
| System Optimization | |||
| Formal | |||
| Games |
| Category | Concrete benchmark members and observed associations |
|---|---|
| Systems & SE (5/11) | BrowseComp, FrontierScience-Olympiad*, HLE, IMO-AnswerBench, SimpleQA, SWE-bench Multilingual, Terminal-Bench, Toolathlon (+0.70); AIME 2026, DeepSWE, SciCode (+0.80); HMMT Feb 2026, HMMT Nov 2025 (+0.82); MathArena, -Banking (+0.90); CritPt, FrontierScience-Research, GPQA (+1.00) |
| Optimization (5/14) | DeepSWE, DeepSearchQA (+0.72); APEX Agents, CyberGym, MCPAtlas (+0.82); Claw Eval, NL2Repo, SWE-bench Pro (+0.97) |
| Knowledge (5/4) | AA-LCR (-0.74); APEX Agents, CyberGym, MCPAtlas (+0.74); Claw Eval, NL2Repo, SWE-bench Pro (+0.95) |
| System Optimization (4/12) | MathArena, MMLU-Pro (-1.00); GPQA, SciCode, SimpleQA (-0.80); AA-LCR, HMMT Nov 2025* (-0.74); IMO-AnswerBench (+1.00) |
| Category | Benchmark | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Last 1 | Last 2 | Last 3 | Last 1 | Last 2 | Last 3 | Last 1 | Last 2 | Last 3 | ||
| CUDA | SWE Pro | |||||||||
| Model Dev. | MathArena | |||||||||
| Puzzle | SciCode | |||||||||
| System Opt. | MMLU-Pro | |||||||||
| System Opt. | AA-LCR | |||||||||
| Category | Concrete benchmark members and observed associations |
|---|---|
| Formal (5/8) | -Banking (+0.70); AA-LCR (+0.80) points: CritPt, FrontierScience-Research, FrontierScience-Olympiad*, GPQA (+0.87); HMMT Feb 2026 (+0.95); AIME 2026, SciCode, -Banking (+0.97) |
| Games (5/8) | HMMT Nov 2025 (+0.71); AA-LCR, AIME 2026, SciCode (+0.72); HMMT Feb 2026 (+0.76); BrowseComp, IMO-AnswerBench, MathArena, SimpleQA (+0.82); -Banking (+0.87); CritPt, FrontierScience-Research, GPQA (+0.97) |
| Scientific & ML (4/4) | APEX Agents, BrowseComp, CritPt, CyberGym, DeepSearchQA, FrontierScience-Research, GPQA, HLE, MMLU-Pro, SimpleQA, SWE-bench Multilingual, SWE-bench Verified, Terminal-Bench, Toolathlon (+0.80); HMMT Nov 2025 (+0.95); DeepSWE, MathArena (+1.00) |
| CUDA (4/3) | HMMT Feb 2026, MathArena (-0.95); BrowseComp*, FrontierScience-Olympiad, IFEval*, SimpleQA (-0.74); MCPAtlas, SWE-bench Multilingual, SWE-bench Pro (+0.74) |
| Model Development (4/2) | FrontierScience-Research, HMMT Feb 2026, SimpleQA (-1.00); BrowseComp*, MathArena (-0.80); SWE-bench Multilingual (+0.80) points: BrowseComp*, FrontierScience-Research, HMMT Feb 2026, SimpleQA (-0.95); MCPAtlas, SWE-bench Multilingual, SWE-bench Pro (+0.74) |
| Puzzle & Challenge (4/5) | AIME 2026*, CritPt, FrontierScience-Olympiad*, GPQA, HMMT Feb 2026*, MathArena*, MMLU-Pro*, SciCode, SimpleQA* (-0.95) |
| Model | First score | Final score | Gain by 60% | |
|---|---|---|---|---|
| Opus 4.7 | 63.96 | 68.55 | 94.8% | 0.179 |
| DeepSeek V4 Pro | 31.40 | 50.92 | 91.0% | 0.255 |
| Gemini 3.1 Pro | 51.13 | 60.45 | 100.0% | 0.125 |
| GLM 5.2 | 50.63 | 63.52 | 71.5% | 0.313 |
| Policy | AutoLab | EdgeBench |
|---|---|---|
| Fixed budget | ||
| Time-based patience | ||
| Recent gain (1 interval) | ||
| Recent gain (2 intervals) | ||
| Conditioned continuation |
| AutoLab | EdgeBench | |||||||
|---|---|---|---|---|---|---|---|---|
| Score retained | Time saved | Score retained | Time saved | |||||
| Method | ||||||||
| Fixed budget | ||||||||
| Time-based patience | ||||||||
| Recent gain (1 interval) | ||||||||
| Recent gain (2 intervals) | ||||||||
| Target | Policy | AutoLab | EdgeBench | ||
|---|---|---|---|---|---|
| Cost | Loss (pt; %) | Cost | Loss (pt; %) | ||
| Fixed budget | ( ) | ( ) | |||
| Conditioned | ( ) | ( ) | |||
| Fixed budget | ( ) | ( ) | |||
| Conditioned | ( ) | ( ) | |||
| Fixed budget | ( ) | ( ) | |||
| Measurement | Reference or source |
|---|---|
| AA-LCR | Artificial Analysis methodology |
| AIME 2026 | MathArena dataset |
| APEX-Agents | Vidgen et al. [2026] |
| BrowseComp | OpenAI evaluation release |
| Claw Eval (pass 3 ) | Claw-Eval release |
| CritPt | CritPt dataset |