LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood. In this paper, we systematically investigate two principal test-time scaling axes of LLM agents: sequential scaling through extended interaction and parallel scaling through trajectory sampling. We first introduce a realistic benchmark that provides one unified framework for evaluating LLM agents across search, coding, reasoning, and tool-use domains, more faithfully reflecting the heterogeneity of real-world deployments. Evaluating ten leading LLM agents reveals substantial performance degradation when transitioning from domain-specific evaluations to this realistic setting. Building on this foundation, we progressively scale test-time compute along fine-grained increments to characterize the performance upper bound. We find that neither scaling axis can consistently yield meaningful gains from additional test-time compute in realistic environments, a phenomenon we attribute to two fundamental limitations: the scaling plateau that bottlenecks sequential scaling and the verification gap that undermines parallel scaling. Code is publicly available at https://github.com/cxcscmu/General-AgentBench.
Figures & tables
Figure 1 : Main contribution of this paper. A : GPT-5’s performance drop under General AgentBench compared to domain-specific evaluation. B : Sequential test-time scaling via longer interaction turns out to be unstable or to degrade performance. C : Agents fail to utilize the improvement brought by multiple sampling due to verification gap.
Domain
Dataset
Original
Sampled
Search
BrowseComp
1266
124
WebVoyager
643
65
Coding
SWE-Bench Verified
500
50
Terminal-Bench
230
80
Reason
MathHay
602
75
Tool-Calling
Tau2-Bench
278
50
Table 1: Composition of General AgentBench .
Figure 2 : Overview of the General AgentBench framework. Green indicates the active task; orange denotes idle but available servers; red indicates excluded data.
Models
Search
Code
Reasoning
Tool-Call
Avg
BrowseComp
WebVoyager
SWE-Bench
Terminal-Bench
MathHay
Tau2-Bench
MCP-Bench
Open-Source
GPT-OSS -120B
8.0 ± 0.17
35.9 ± 0.64
15.0 ± 0.16
18.1 ± 1.26
46.7 ± 0.56
63.0 ± 5.12
67.3 ± 1.63
36.6
4.0 ± 0.25
27.7 ± 0.84
12.0 ± 0.62
6.3 ± 0.35
38.7 ± 0.49
26.0 ± 1.32
63.3 ± 2.92
26.1
▼ 50.0%
▼ 22.8%
▼ 20.0%
▼ 65.2%
▼ 17.1%
▼ 58.7%
▼ 5.9%
▼ 28.7%
Qwen3-235B A22B
12.0 ± 0.28
40.2 ± 2.55
30.0 ± 0.68
40.1 ± 0.78
45.3 ± 0.49
50.0 ± 5.05
74.7 ± 2.55
41.5
Table 2 : Per-benchmark decomposition of General AgentBench . For each model we report performance (mean ± std over n=3 runs) on every constituent benchmark, grouped by domain; the first row is the Baseline ( B ) setting, the second row is the General ( G ) setting, and the third row shows the relative change ( Δ% ) from B to G : red ▼ denotes degradation, green ▲ denotes improvement. Bold marks the best and underline the second-best score.
Figure 3 : Sequential scaling behavior of 5 models over 4 domains on General AgentBench. The x -axis is the context budget and the y -axis is accuracy. The shaded area represents the bootstrapped 95% CI. Performance generally improves at first but saturates or degrades beyond a model- and domain-dependent scaling plateau, sometime even degrades at first, indicating an inherent limitation.
Figure 4 : Instance-level correctness dynamics. We randomly sample 10 instances from the General AgentBench and track their correctness across context budgets. Red indicates incorrect cases and green indicates correct cases. The bottom “ideal case” illustrates how we expect the model to behave.
Figure 5 : Empirical proof of the scaling plateau definition on Qwen3-235B in the Search domain.
Figure 6 : Ablation studies on sequential scaling strategies.
Figure 7 : Verification gap between generation and self-choice. Across four domains, we observe a consistent gap between solution generation and verification: as the number of samples increases, correct solutions appear more frequently in the sampled set, yet models often fail to identify and select them. The dashed and dotted curves represent two self-choice strategies, while the diamond denotes a stronger evaluator, GPT-5.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Accuracy curves and confidence lower bounds Ck for all model–domain pairs. Each panel shows the estimated accuracy μ^k (line) and the one-sided lower bound Ck (bars) at each context-budget transition. Green bars indicate Ck>0 (significant gain); red bars indicate Ck≤0 . The dashed vertical line marks the identified scaling plateau ck∗ .
Figure 9 : Paired-bootstrap distributions for all model–domain pairs. Each violin depicts the resampled distribution of the marginal gain dˉ(k)∗ across B=10,000 bootstrap iterations. The diamond marks the observed dˉ(k) , and the red tick indicates the FPC-corrected lower bound Ck . Blue violins: transitions before the plateau; pink violins: Ck≤0 .
Benchmark
N
Typical CI Width
Pass Std
SWE-Bench
49–50
0.04–0.08
0.02–0.05
Tau2-Bench
50
∼ 0.18
0.02–0.04
Terminal-Bench
79
0.01–0.09
0.01–0.03
MathHay
75
0.11–0.14
0.01–0.03
BrowseComp
124
0.10–0.12
0.01–0.02
MCP-Bench
52
0.02–0.07
0.01–0.03
Appendix
Table 3: Bootstrap statistics across benchmarks. N denotes the number of sampled instances. Typical CI width and bootstrap standard error (Pass Std) are computed over 10,000 resamples across all evaluated models.
Task Category
Sub-category
Representative Benchmarks / Papers
Long Document QA
Single/Multi-Doc Understanding
NarrativeQA Kočiskỳ et al. [2018] , 2WikiMultihopQA Ho et al. [2020] , Loong Wang et al. [2024b]
Long Dialogue Understanding
MeetingBank Hu et al. [2023] , CharacterChat Tu et al. [2023]
Code Understanding
RepoBench Liu et al. [2023a] , CrossCodeEval Ding et al. [2023]
Structured Data Understanding
TabFact Chen et al. [2019] , WikiTableQuestions Kweon et al. [2023] , LongBench Bai et al. [2024] (L-Data)
Summarization
Global Information Aggregation
GovReport Huang et al. [2021] , Multi-News Fabbri et al. [2019] , ZeroSCROLLS Shaham et al. [2023]
Many-Shot
In-Context Learning
MIR-Bench Yan et al. [2025] , Many-Shot ICL Agarwal et al. [2024]
Appendix
Table 4: Taxonomy of Existing Long-Context Benchmarks.
Models
LongBench
HELMET
MRCR
Avg
Open-Source
GPT-OSS-120B
47.8
12.9
32.8
31.2
Qwen3-235B-A22B
58.3
40.8
40.6
46.6
Qwen3-Next
53.1
26.3
27.9
35.8
DeepSeek-V3.2
50.3
48.0
33.2
43.8
DeepSeek-R1
58.3
36.9
39.2
44.8
Appendix
Table 5 : Performance on long-context benchmarks. We report scores across three established benchmarks focusing on retrieval and single-turn comprehension. Bold marks the best and underline the second-best score in each column.
Figure 10 : Pairwise correlation between static long-context benchmarks and agentic domains in our General AgentBench. “All” denotes the average performance across the search, code, reasoning, and tool-use domains.
Figure 11 : Extended parallel scaling to K=8 and K=16 . We evaluate Gemini-2.5-Flash and DeepSeek-V3.2 on coding and tool-call domains with three commitment strategies and GPT-5 as an external verifier. The verification gap persists across all K values; majority voting growth decelerates while pairwise selection shows an accelerating trend.
Verifier
Run 1
Run 2
Run 3
Std
Self-judge (pointwise)
0.1795
0.1823
0.1705
0.0050
Self-judge (pairwise)
0.1795
0.1795
0.1672
0.0058
GPT-5 (pointwise)
0.1397
0.1471
0.1397
0.0035
Appendix
Table 6: Verification accuracy across three independent runs (Gemini-2.5-Flash, coding domain). Low standard deviations confirm that both self-judgment and external verification are stable across runs.
Benchmark
Out-of-domain tool calls (%)
SWE-Bench
∼ 5
Terminal-Bench
∼ 9
MCP-Bench
∼ 8
BrowseComp
∼ 13
Tau2-Bench
∼ 40
WebVoyager
∼ 50
Appendix
Table 7: Proportion of out-of-domain tool calls per benchmark under the unified setting (Gemini-2.5-Flash).
Tool registry size
24K
32K
36K
42K (Original)
Gemini-2.5-Flash
0.3550
0.3594
0.3572
0.3478
Appendix
Table 8: Performance of Gemini-2.5-Flash on the tool-call domain as the tool registry size varies.
Model
Input
Cached Input
Output
Gemini-2.5-Flash
0.30
0.03
2.50
Gemini-2.5-Pro
1.25
0.125
10.00
GPT-5
1.25
0.125
10.00
Claude-Haiku 4.5
1.00
0.50
5.00
Claude-Sonnet 4.5
3.00
3.75
15.00
gpt-oss-120B
0.15
0.075
0.60
Appendix
Table 9: Unit API prices (USD per 1M tokens) used in our cost estimation.
Model
Search
MathHay
SWEBench
MCPBench
Tau2Bench
TerminalBench
Total
Gemini-2.5-Flash
$193
$70.9
$12719
$122
$62.2
$826
$13993
DeepSeek-R1
$5188
$157
$3675
$492
$202
$1505
$11218
DeepSeek-V3.2
$369
$38.8
$7.80
$88.4
$54.4
$412
$970
Qwen3-235B
$1254
$41.4
$1286
$120
$55.5
$409
$3166
Qwen3-Next
$25.3
$9.52
$3.38
$21.2
$14.9
$154
$229
Total
$7028
$317
$17692
$843
$389
$3307
$29576
Appendix
Table 10 : Cost for evaluating models under parallel scaling setting (USD)
Model
Search
MathHay
SWEBench
MCPBench
Tau2Bench
TerminalBench
Total
Gemini-2.5-Flash
$5588
$369
$2024
$568
$852
$1870
$11271
DeepSeek-R1
$1267
$654
$892
$931
$1088
$782
$5614
DeepSeek-V3.2
$902
$333
$191
$79.1
$52.3
$356
$1913
Qwen3-235B
$1370
$238
$436
$716
$162
$689
$3610
Qwen3-Next
$442
$219
$392
$152
$251
$529
$1985
Total
$9568
$1814
$3935
$2445
$2405
$4225
$24392
Appendix
Table 11 : Cost for evaluating models under sequential scaling setting (USD)
Model
Search
MathHay
SWEBench
MCPBench
Tau2Bench
TerminalBench
Total
Gemini-2.5-Pro
$47.5
$18.6
$3253
$31.3
$15.2
$203
$3569
GPT-5
$87.5
$14.2
$146
$21.0
$12.4
$60.9
$342
Claude-Haiku-4.5
$304
$10.9
$312
$29.3
$14.1
$106
$776
Claude-Sonnet-4.5
$1248
$40.9
$926
$129
$52.0
$376
$2772
OpenAI-oss-120B
$1.44
$1.49
$0.52
$0.50
$1.35
$0.40
$5.70
DeepSeek-V3.2
$96.8
$10.0
$1.92
$22.0
$14.0
$102
$247
Appendix
Table 12 : Cost for evaluating models under general (default context) setting (USD)
Context
DeepSeek-R1
Qwen3-Next
Qwen3-235B
DeepSeek-V3.2
Gemini2.5-Flash
Search
64k
0.2060
0.2186
0.0510
0.0854
0.1558
80k
0.1210
0.2111
0.1256
0.2663
0.1658
96k
0.0550
0.2312
0.1457
0.2410
0.1759
112k
0.1260
0.2663
0.2010
0.2320
0.1759
128k
0.1260
0.2500
0.1809
0.2320
0.1759
Appendix
Table 13 : Detailed numerical results for sequential scaling (Figure 3 ). Accuracy of 5 models across 4 domains at increasing context budgets.
Context
Qwen3-235B
Gemini-2.5-Flash
Memory
Verifier
Ours
Memory
Verifier
Ours
Tool-call
64k
0.446
0.473
0.458
0.477
0.504
0.489
80k
0.462
0.491
0.474
0.326
0.454
0.338
96k
0.479
0.489
0.471
0.317
0.446
0.329
128k
0.449
0.488
0.441
0.317
0.387
0.329
Appendix
Table 14: Detailed numerical results for the ablation on sequential scaling strategies (Figure 6 ). Three protocols—Memory (summarization), Verifier (verifier-guided revision), and Ours (naive scaling)—are compared on two models across two domains at five context budgets.
Model
K
Pass@ K
Majority
Self-Choice
Pairwise
GPT5
DeepSeek-R1
1
0.2010
0.2010
0.2010
0.2010
0.0704
2
0.2460
0.2202
0.2111
0.2111
0.0754
3
0.2760
0.2387
0.2161
0.2161
0.0955
4
0.2970
0.2286
0.2211
0.2211
0.1156
Qwen3-Next
1
0.1859
0.1859
0.1859
0.1859
0.0352
2
0.2462
0.2167
0.1357
0.2010
0.0653
Appendix
Table 15: Verification gap – Search (Figure 7 ). Pass@ K (oracle upper bound) and four commitment strategies at K=1,2,3,4 .
Model
K
Pass@ K
Majority
Self-Choice
Pairwise
GPT5
DeepSeek-R1
1
0.1076
0.1076
0.1076
0.1076
0.1084
2
0.1690
0.1520
0.1548
0.1640
0.1626
3
0.2462
0.2312
0.2249
0.2265
0.2094
4
0.2620
0.2380
0.2483
0.2343
0.2249
Qwen3-Next
1
0.0923
0.0923
0.0923
0.0923
0.0929
2
0.1385
0.1335
0.1162
0.1239
0.1239
Appendix
Table 16: Verification gap – Code (Figure 7 ).
Model
K
Pass@ K
Majority
Self-Choice
Pairwise
GPT5
DeepSeek-R1
1
0.4667
0.4667
0.4667
0.4667
0.4267
2
0.5600
0.4932
0.4267
0.4800
0.5067
3
0.5867
0.4824
0.4933
0.4533
0.4933
4
0.6133
0.4954
0.5200
0.4400
0.5178
Qwen3-Next
1
0.4200
0.4200
0.4200
0.4200
0.4000
2
0.6000
0.5368
0.4667
0.5330
0.4733
Appendix
Table 17: Verification gap – Reason (Figure 7 ).
Model
K
Pass@ K
Majority
Self-Choice
Pairwise
GPT5
DeepSeek-R1
1
0.3968
0.3968
0.3968
0.3968
0.2200
2
0.4716
0.4611
0.4284
0.4446
0.2800
3
0.4873
0.4621
0.4297
0.4470
0.2800
4
0.5331
0.5061
0.4748
0.4912
0.3050
Qwen3-Next
1
0.5678
0.5678
0.5678
0.5678
0.1800
2
0.6418
0.5969
0.5775
0.5842
0.2800
Appendix
Table 18: Verification gap – Tool-use (Figure 7 ).
Figure 12 : Attention behavior of Qwen3-235B (full attention) under the General AgentBench setting.
Figure 13 : Attention behavior of Qwen3-Next (hybrid linear attention) under the General AgentBench setting.
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submissions, making progress observable throughout long trajectories. We propose Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales. We apply it to four general-purpose agents on four open-ended benchmarks, with sessions of up to 100M tokens, and to three feedback-driven LLM optimization harnesses in controlled single-task interventions. Independent sampling provides a theoretically characterized reference, for which Elo grows linearly with log compute. Against this reference, agents can initially convert tokens into Elo faster than independent sampling, but their marginal gains diminish and eventually fall below the reference. In contrast, the strongest historical human contestants improve superlinearly over contest time on shared AtCoder Heuristic Contest tasks, providing evidence of continual learning and substantial headroom after agents slow down. We define the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference. Using this point as the per-session budget, we split 100M tokens across parallel sessions on FrontierCS Polyomino Packing, gaining +264 Elo over one long session and +355 over ten short sessions.
Kaiyuan Liu, Qiuyang Mang, Bo Peng +6
UC Berkeley · University of Washington · Princeton University +1
Test-time scaling has become a powerful way to improve large language models. However, existing methods are best suited to short, bounded outputs that can be directly compared, ranked or refined. Long-horizon coding agents violate this premise: each attempt produces an extended trajectory of actions, observations, errors, and partial progress taken by the agent. In this setting, the main challenge is no longer generating more attempts, but representing prior experience in a form that can be effectively selected from and reused. We propose a test-time scaling framework for agentic coding based on compact representations of rollout trajectories. Our framework converts each rollout into a structured summary that preserves its salient hypotheses, progress, and failure modes while discarding low-signal trace details. This representation enables two complementary forms of inference-time scaling. For parallel scaling, we introduce Recursive Tournament Voting (RTV), which recursively narrows a population of rollout summaries through small-group comparisons. For sequential scaling, we adapt Parallel-Distill-Refine (PDR) to the agentic setting by conditioning new rollouts on summaries distilled from prior attempts. Our method consistently improves the performance of frontier coding agents across SWE-Bench Verified and Terminal-Bench v2.0. For example, by using our method Claude-4.5-Opus improves from 70.9% to 77.6% on SWE-Bench Verified (mini-SWE-agent) and 46.9% to 59.1% on Terminal-Bench v2.0 (Terminus 1). Our results suggest that test-time scaling for long-horizon agents is fundamentally a problem of representation, selection, and reuse.
Joongwon Kim, Wannan Yang, Kelvin Niu +13
1Meta Superintelligence Labs · University of Washington · 3New York University +4
Test-time scaling (TTS) has become an effective approach for improving large language model performance by allocating additional computation during inference. However, existing TTS strategies are largely hand-crafted: researchers manually design reasoning patterns and tune heuristics by intuition, leaving much of the computation-allocation space unexplored. We propose an environment-driven framework, AutoTTS, that changes what researchers design: from individual TTS heuristics to environments where TTS strategies can be discovered automatically. The key to AutoTTS lies in environment construction: the discovery environment must make the control space tractable and provide cheap, frequent feedback for TTS search. As a concrete instantiation, we formulate width--depth TTS as controller synthesis over pre-collected reasoning trajectories and probe signals, where controllers decide when to branch, continue, probe, prune, or stop and can be evaluated cheaply without repeated LLM calls. We further introduce beta parameterization to make the search tractable and fine-grained execution trace feedback to improve discovery efficiency by helping the agent diagnose why a TTS program fails. Experiments on mathematical reasoning benchmarks show that the discovered strategies improve the overall accuracy--cost tradeoff over strong manually designed baselines. The discovered strategies generalize to held-out benchmarks and model scales, while the entire discovery costs only $39.9 and 160 minutes. Our data, and code will be open-source at https://github.com/zhengkid/AutoTTS.