Long-horizon agents improve solutions through sustained interaction, execution, and task feedback. Scaling studies relate performance to resources and capabilities, yet how existing capabilities shape returns to extended interaction remains less understood. To address this gap, we analyze AutoLab and EdgeBench, two long-horizon benchmarks. We find that starting performance and subsequent growth are associated with different capabilities: within a task category, similar early scores can precede different later gains. To formalize this finding, we model capability-time scaling with category-specific logistic power laws shared across models. Fitted to early trajectories, these curves extrapolate the observed models' category-average scores to later computation. However, rising average scores mask narrowing improvement opportunities: later gains concentrate among fewer improving models. High final scores and continued improvement also have distinct capability profiles. Predicted mean gains estimate each model's fraction of improving tasks; averaging these estimates forecasts the average share of improving models. These uneven returns motivate deciding whether a specific run should continue. We therefore derive a continuation policy to save time and compute with limited score loss. The policy conditions growth predictions on the run's observed progress and weighs immediate and delayed gains against computation costs. In replay with training and price calibration based on other models' histories, the policy saves roughly one-third of full-run time, with relative score losses of 2.4% on AutoLab individual runs and 3.3% on EdgeBench published mean curves. Our repository is available at https://github.com/Chihaya-Anon-chan/long-horizon-scaling.
Figures & tables
Figure 1: From capability–time scaling to adaptive continuation. (1) All 254 observations versus the fitted index. The gray curve shows the logistic link; filled/open markers denote early/later observations. RMSE is averaged across categories. (2) Thin lines show ten categories; bold lines show their means. Scores use first/last interval endpoints (AutoLab 40/100% budget; EdgeBench 4/12 h). (3) Growth fitted on other models and the current prefix guide stopping. Reported tradeoffs use prices calibrated on other models and then fixed (AutoLab/EdgeBench; Section 5.3 ).
Figure 2: Shared capabilities organize category–time responses across all ten categories: AutoLab (a–d), EdgeBench (e–j). Filled/open markers show early/later category means, with model identity fixed by color and shape; all 254 means are retained. Gray curves show 100σ(η) , and annotations report later R2 and RMSE in score points. Section 3.2 specifies the early/later checkpoints and five-PC fitting protocol; Table 7 lists the category panels. Axis ranges vary by category.
Figure 3: Observed trajectories show early gain realization and sustained growth. Fits from Figure 2 ; all 50 model–checkpoint means: (a) 12 tasks, four models; (b) eight tasks, five models. Filled/solid: early observations/fits; open/dashed: later observations/predictions. Dotted: fitting cutoffs.
Figure 4: Fewer models improve later, with category-specific capability profiles. (a) First/last-interval participation and gain coverage across ten categories. (b) Task-weighted group means of predicted and observed improvement frequencies for all 90 later observations. Groups use early prediction quantiles; bars show 95% whole-task bootstrap intervals. Dashed lines indicate equality. (c) Systems & SE: benchmark correlations with final score and terminal improvement frequency. (d) Optimization’s eight benchmarks with ρs≥0.70 for terminal frequency, compared with Knowledge. Correlations use five models per category; terminal frequency averages the last two intervals. Lines pair the same benchmark. Details and measurements: Appendix B .
Figure 5: Conditioned continuation improves cost–loss tradeoffs. (a–b) Frontiers averaged over replicates then models. (c) Reductions in mean score loss versus fixed-budget stopping; all seven AutoLab and five EdgeBench evaluations improve. Comparisons use the same cost range across policies for each model and replicate; (a–b) display the range shared across each suite’s evaluations.
AutoLab
EdgeBench
Matched condition
Reported outcome
Fixed budget
Ours
Fixed budget
Ours
Save 30% time
Score retained (%)
89.9
97.3
94.7
96.8
Save 50% time
Score retained (%)
77.1
90.5
89.5
91.8
Retain 95% score
Time saved (%)
17.2
38.1
28.5
38.4
Retain 90% score
Time saved (%)
29.7
50.9
48.2
55.9
Table 1: Matched operating points, interpolated from the mean frontiers in Figure 5 . Score retention is relative to the mean full-run score. Ours denotes conditioned continuation; higher is better. Table 14 compares all baselines.
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Number of PCs
1
3
5
7
29
Mean later-score RMSE
4.861
3.180
2.315
2.288
2.504
Appendix
Table 2: Dimension sweep and early-only dimension selection on the same temporal panels. Metrics average the ten categories equally; RMSE is in score points.
Category
Affine in log time
Logistic power law
CUDA
4.42
4.02
Model Development
5.68
5.73
Puzzle & Challenge
5.86
4.28
System Optimization
3.91
3.48
Formal
2.03
1.53
Games
1.09
1.07
Appendix
Table 3: Matched link-family comparison using all 29 external inputs. Both families select regularization and slope structure on early checkpoints and predict the same 90 later means. Errors are RMSE in score points.
Figure 6: From external capabilities to time responses. (a) The union of the two strongest positive and negative loadings per PC shows 11 benchmarks; PC1 has only positive loadings. All 29 benchmarks enter the representation; full loadings appear in Figure 9 . (b–c) Illustrative categories: top, fitted PC coefficients wc,vc for A,B (intercepts omitted); bottom, all 50 observed category-mean scores across four/five models on 12/8 fixed tasks and their fitted curves. Filled/solid denotes early observations/fits; open/dashed denotes later observations/predictions. Both use the five-PC response of Figure 2 . Loadings and readout coefficients use separate color scales.
Benchmark
Early coefficient fit
Later evaluation
AutoLab
20%,40%,60%
80%,100%
EdgeBench
2,4,6,8 h
10,12 h
Appendix
Table 4: Temporal prediction protocol. AutoLab time is a fraction of declared budget; EdgeBench time is elapsed hours.
Category
Last score
Independent
Indicators
PCs
CUDA
4.19
4.95
4.18
3.75
Model Development
6.19
3.46
6.06
3.79
Puzzle & Challenge
5.83
4.27
4.41
4.08
System Optimization
1.71
3.97
3.41
3.46
Formal
4.40
1.86
1.66
1.93
Games
1.95
1.10
1.01
1.03
Appendix
Table 5: Later-score RMSE in points under matched temporal evaluation. Indicators denotes regularized model identities; PCs denotes the selected five-PC response.
Fixed primary panel
Expanded panel
Category
Tasks
Decline (pp)
Tasks
Decline (pp)
CUDA
3
41.7
4
32.5
Model Development
2
50.0
2
42.5
Puzzle & Challenge
5
30.0
9
43.0
System Optimization
12
54.2
14
48.3
Formal
8
20.0
8
20.0
Appendix
Table 6: Participation declines under both panel definitions. Counts are tasks; decline is the task-averaged first-to-last reduction in percentage points. Expanded panels retain every available complete task–model curve, with each task's model set fixed across intervals.
Figure 7: The gain–opportunity response and later calibration. (a) All 209 model–category–interval observations (filled: early; open: later). The black curve fits early improvement frequencies; orange squares connect later averages in five groups defined by early gain-ratio quantiles. (b) All 90 later fractions (gray), with the same task-weighted means and 95% whole-task bootstrap intervals as Figure 4 (b); dashed line: predicted = observed.
Category
Models/tasks
Task process
CUDA
4/3
Optimize decoding, geometric correspondence, and elliptic-curve CUDA kernels.
Model Development
4/2
Select fine-tuning data and optimize an online serving engine.
Puzzle & Challenge
4/5
Reduce program or model complexity under correctness and accuracy constraints.
System Optimization
4/12
Improve implementations of cryptography, retrieval, storage, and numerical kernels.
Formal
5/8
Complete interdependent Lean/Coq proofs in multi-file projects.
Games
5/8
Explore interactive fiction or implement game-playing and management agents.
Appendix
Table 7: Work performed by the fixed task panels. Counts are models/tasks. The first four categories are AutoLab; the remaining six are EdgeBench. The machine-readable supplement retains all 71 task descriptions.
Category
Final score
Terminal frequency
Difference
CUDA
0.290
−0.222
+0.511
Model Development
0.250
−0.250
+0.500
Puzzle & Challenge
0.473
−0.364
+0.837
System Optimization
0.304
−0.051
+0.354
Formal
0.671
0.171
+0.499
Games
0.444
0.601
−0.156
Appendix
Table 8: Attainment and terminal improvement have different external benchmark profiles. Each entry averages signed model-rank correlations over all 29 measurements within a category; the final row averages categories equally. Difference is final-score minus terminal-frequency association.
Category
Concrete benchmark members and observed associations
Table 9: Composite profiles in the main-text categories. All external measurements with ∣ρs∣≥0.70 for terminal improvement frequency are listed, with no cap on profile size. Parentheses give Spearman correlations; category counts are models/tasks. The default uses the last two intervals and positive gain; the additional increment threshold is stated explicitly. An asterisk identifies a benchmark with at least one frozen estimated input in that panel. Correlated members describe a joint performance profile, not separate effects.
Category
Benchmark
100Δ>0
100Δ>0.1
100Δ>0.5
Last 1
Last 2
Last 3
Last 1
Last 2
Last 3
Last 1
Last 2
Last 3
CUDA
SWE Pro
+0.89
+0.74
+0.74
+0.74
+0.60
+0.60
+0.74
+0.60
+0.60
Model Dev.
MathArena
−0.45
−0.80
−0.80
−0.45
−0.80
−0.80
−0.45
−0.63
−0.40
Puzzle
SciCode
−0.89
−0.95
−0.95
−0.89
−0.95
−0.95
−0.89
−0.95
−0.77
System Opt.
MMLU-Pro
−1.00
−1.00
−1.00
−1.00
−1.00
−1.00
−0.95
−1.00
−1.00
System Opt.
AA-LCR
−0.74
−0.74
−0.74
−0.74
−0.74
−0.74
−0.89
−0.74
−0.74
Appendix
Table 10: Window and gain-threshold sensitivity of the representative associations. Entries are Spearman correlations across fixed model panels with the fraction of tasks improving, averaged over the indicated final intervals. Last two is the main window; thresholds are in 0–100 score points. The first five rows concern AutoLab, the remainder EdgeBench. Every displayed external measurement is published for all models in its panel.
Category
Concrete benchmark members and observed associations
Table 11: Composite profiles in the remaining six categories. All external measurements with ∣ρs∣≥0.70 for terminal improvement frequency are listed, with no cap on profile size. Parentheses give Spearman correlations; category counts are models/tasks. The default uses the last two intervals and positive gain; the additional increment threshold is stated explicitly. An asterisk identifies a benchmark with at least one frozen estimated input in that panel. Correlated members describe a joint performance profile, not separate effects.
Figure 8: Window sensitivity on all ten categories. (a) Benchmark correlations with the equally weighted mean fraction of tasks improving across the selected final intervals, using positive gain. The displayed pairs come from the task-interpretation analysis. (b) Task-averaged model participation decline relative to the first interval, in percentage points. The horizontal rule separates AutoLab from EdgeBench; model/task counts follow Table 7 . Last two is the main window. All tasks and models are retained, and the response models are unchanged.
Model
First score
Final score
Gain by 60%
ℓgain
Opus 4.7
63.96
68.55
94.8%
0.179
DeepSeek V4 Pro
31.40
50.92
91.0%
0.255
Gemini 3.1 Pro
51.13
60.45
100.0%
0.125
GLM 5.2
50.63
63.52
71.5%
0.313
Appendix
Table 12: Gain realization on all four models in AutoLab System Optimization (12 tasks). Scores use the 0–100 scale. The fourth column is the fraction of each model's observed 20–100% gain realized by 60% budget. The timing index ℓgain weights normalized interval midpoints by category mean-score increments; smaller values indicate earlier gains.
Policy
AutoLab
EdgeBench
Fixed budget
14.384
2.442
Time-based patience
19.372
3.233
Recent gain (1 interval)
22.907
1.946
Recent gain (2 intervals)
24.271
1.982
Conditioned continuation
8.825
1.774
Appendix
Table 13: Mean score loss over matched cost ranges (score points; lower is better). Ranges are shared across policies for each model and replicate; results average replicates then models.
AutoLab
EdgeBench
Score retained
Time saved
Score retained
Time saved
Method
30%
50%
95%
90%
30%
50%
95%
90%
Fixed budget
89.9
77.1
17.2
29.7
94.7
89.5
28.5
48.2
Time-based patience
81.3
65.9
8.2
18.0
92.8
84.9
24.6
37.2
Recent gain (1 interval)
75.1
58.6
6.0
12.1
96.2
91.6
38.1
54.9
Recent gain (2 intervals)
72.6
54.7
5.5
10.9
96.2
91.8
34.9
54.6
Appendix
Table 14: Matched tradeoffs on the mean frontiers in Figure 5 . For each suite: full-run score retained at 30%/50% time savings, and time saved at 95%/90% score retention. All entries are percentages; higher is better, best at displayed precision in bold. Values use interpolation within shared cost ranges.
Target
Policy
AutoLab
EdgeBench
Cost
Loss (pt; %)
Cost
Loss (pt; %)
25%
Fixed budget
46.3%
15.77 ( 24.73% )
16.7%
10.12 ( 28.36% )
Conditioned
46.5%
6.99 ( 10.96% )
23.3%
7.66 ( 21.46% )
37.5%
Fixed budget
46.3%
15.77 ( 24.73% )
33.3%
5.95 ( 16.68% )
Conditioned
67.2%
1.52 ( 2.39% )
35.5%
4.96 ( 13.90% )
50%
Fixed budget
70.1%
5.34 ( 8.38% )
50.0%
3.74 ( 10.49% )
Appendix
Table 15: Frozen operating points. Each target supplies the conditioned policy’s source budget constraint and the fixed policy’s declared cutoff. Cost is a fraction of full-run time. Loss is in points, with relative score loss in parentheses. Results average replicates within model, then models.
Figure 9: External representation used by the capability–time model. Left: all 29 benchmark loadings on five PCs estimated from 82 reference model entries, excluding evaluated versions and configuration aliases. Right: the fixed coordinates of all ten evaluated model versions. Loadings and coordinates have separate color scales. The loadings summarize covariance patterns among external measurements; missing-entry estimation is specified in Appendix A .
Measurement
Reference or source
AA-LCR
Artificial Analysis methodology
AIME 2026
MathArena dataset
APEX-Agents
Vidgen et al. [2026]
BrowseComp
OpenAI evaluation release
Claw Eval (pass 3 )
Claw-Eval release
CritPt
CritPt dataset
Appendix
Table 16: Sources of the 29 external measurements. Rows identify the versions or metrics used; links point to benchmark papers, dataset releases, or evaluation documentation.
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understand the current state of this capability, however, evaluation must go beyond final scores, which neither reveal where progress is gained or lost nor indicate whether accumulated experience improves later decisions. We therefore present a systematic evaluation of seven frontier models on 36 long-horizon tasks based on a new framework that uses rule-based metrics to characterize within-run behavior through Solution Framing, Execution, and Feedback Control and controlled comparisons to assess experience reuse within and across tasks. The results show that current agents operate more like engineering optimizers than fully autonomous researchers: they can formulate and implement practical solutions, but their performance varies substantially across runs, their strongest solutions mainly adapt or combine established techniques, and genuine methodological novelty remains rare. Detailed analysis reveals that observed performance is shaped by multiple factors, including distinct process bottlenecks behind similar final outcomes, experience reuse that can help or mislead subsequent decisions, and harness designs that affect performance stability. These findings suggest concrete directions for improving model training, inference-time strategies, experience management, and harness design.
Yiwei Li, Wanli Yang, Hexiang Tan +10
1Meituan · University of Chinese Academy of Sciences
Scientific and engineering progress is fundamentally a long-horizon iterative process: proposing changes, running experiments, measuring outcomes, and continuously refining artifacts. Yet existing benchmarks for frontier models primarily evaluate either single-turn responses or short-horizon agent trajectories, failing to capture the challenges of sustained iterative improvement over extended time horizons. To address this gap, we introduce AutoLab, a new benchmark for ultra long-horizon closed-loop optimization. AutoLab consists of 36 realistic, expert-curated tasks spanning four diverse domains: system optimization, puzzle & challenge, model development, and CUDA kernel optimization. Each task begins with a correct but deliberately suboptimal baseline and challenges agents to improve it within a strict wall-clock budget. Evaluating 17 state-of-the-art models reveals the dominant predictor of success is not the quality of an agent's initial attempt, but its persistence in repeatedly benchmarking, editing, and incorporating empirical feedback. While claude-opus-4.6 exhibits strong long-horizon optimization capabilities, most frontier models, including several proprietary ones, either terminate prematurely or exhaust their budgets with minimal progress. These results underscore the importance of time awareness and persistent iteration in autonomous agents. We open-source the full benchmark, evaluation harness, and task artifacts, to accelerate research toward truly capable long-horizon agents.
AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.
Zongxia Li, Zhongzhi Li, Yucheng Shi +10
Tencent HY LLM Frontier · University of Maryland, College Park · University of Georgia +5