LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood. In this paper, we systematically investigate two principal test-time scaling axes of LLM agents: sequential scaling through extended interaction and parallel scaling through trajectory sampling. We first introduce a realistic benchmark that provides one unified framework for evaluating LLM agents across search, coding, reasoning, and tool-use domains, more faithfully reflecting the heterogeneity of real-world deployments. Evaluating ten leading LLM agents reveals substantial performance degradation when transitioning from domain-specific evaluations to this realistic setting. Building on this foundation, we progressively scale test-time compute along fine-grained increments to characterize the performance upper bound. We find that neither scaling axis can consistently yield meaningful gains from additional test-time compute in realistic environments, a phenomenon we attribute to two fundamental limitations: the scaling plateau that bottlenecks sequential scaling and the verification gap that undermines parallel scaling. Code is publicly available at https://github.com/cxcscmu/General-AgentBench.
Figures & tables
Figure 1 : Main contribution of this paper. A : GPT-5’s performance drop under General AgentBench compared to domain-specific evaluation. B : Sequential test-time scaling via longer interaction turns out to be unstable or to degrade performance. C : Agents fail to utilize the improvement brought by multiple sampling due to verification gap.
Domain
Dataset
Original
Sampled
Search
BrowseComp
1266
124
WebVoyager
643
65
Coding
SWE-Bench Verified
500
50
Terminal-Bench
230
80
Reason
MathHay
602
75
Tool-Calling
Tau2-Bench
278
50
Table 1: Composition of General AgentBench .
Figure 2 : Overview of the General AgentBench framework. Green indicates the active task; orange denotes idle but available servers; red indicates excluded data.
Models
Search
Code
Reasoning
Tool-Call
Avg
BrowseComp
WebVoyager
SWE-Bench
Terminal-Bench
MathHay
Tau2-Bench
MCP-Bench
Open-Source
GPT-OSS -120B
8.0 ± 0.17
35.9 ± 0.64
15.0 ± 0.16
18.1 ± 1.26
46.7 ± 0.56
63.0 ± 5.12
67.3 ± 1.63
36.6
4.0 ± 0.25
27.7 ± 0.84
12.0 ± 0.62
6.3 ± 0.35
38.7 ± 0.49
26.0 ± 1.32
63.3 ± 2.92
26.1
▼ 50.0%
▼ 22.8%
▼ 20.0%
▼ 65.2%
▼ 17.1%
▼ 58.7%
▼ 5.9%
▼ 28.7%
Qwen3-235B A22B
12.0 ± 0.28
40.2 ± 2.55
30.0 ± 0.68
40.1 ± 0.78
45.3 ± 0.49
50.0 ± 5.05
74.7 ± 2.55
41.5
Table 2 : Per-benchmark decomposition of General AgentBench . For each model we report performance (mean ± std over n=3 runs) on every constituent benchmark, grouped by domain; the first row is the Baseline ( B ) setting, the second row is the General ( G ) setting, and the third row shows the relative change ( Δ% ) from B to G : red ▼ denotes degradation, green ▲ denotes improvement. Bold marks the best and underline the second-best score.
Figure 3 : Sequential scaling behavior of 5 models over 4 domains on General AgentBench. The x -axis is the context budget and the y -axis is accuracy. The shaded area represents the bootstrapped 95% CI. Performance generally improves at first but saturates or degrades beyond a model- and domain-dependent scaling plateau, sometime even degrades at first, indicating an inherent limitation.
Figure 4 : Instance-level correctness dynamics. We randomly sample 10 instances from the General AgentBench and track their correctness across context budgets. Red indicates incorrect cases and green indicates correct cases. The bottom “ideal case” illustrates how we expect the model to behave.
Figure 5 : Empirical proof of the scaling plateau definition on Qwen3-235B in the Search domain.
Figure 6 : Ablation studies on sequential scaling strategies.
Figure 7 : Verification gap between generation and self-choice. Across four domains, we observe a consistent gap between solution generation and verification: as the number of samples increases, correct solutions appear more frequently in the sampled set, yet models often fail to identify and select them. The dashed and dotted curves represent two self-choice strategies, while the diamond denotes a stronger evaluator, GPT-5.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Accuracy curves and confidence lower bounds Ck for all model–domain pairs. Each panel shows the estimated accuracy μ^k (line) and the one-sided lower bound Ck (bars) at each context-budget transition. Green bars indicate Ck>0 (significant gain); red bars indicate Ck≤0 . The dashed vertical line marks the identified scaling plateau ck∗ .
Figure 9 : Paired-bootstrap distributions for all model–domain pairs. Each violin depicts the resampled distribution of the marginal gain dˉ(k)∗ across B=10,000 bootstrap iterations. The diamond marks the observed dˉ(k) , and the red tick indicates the FPC-corrected lower bound Ck . Blue violins: transitions before the plateau; pink violins: Ck≤0 .
Benchmark
N
Typical CI Width
Pass Std
SWE-Bench
49–50
0.04–0.08
0.02–0.05
Tau2-Bench
50
∼ 0.18
0.02–0.04
Terminal-Bench
79
0.01–0.09
0.01–0.03
MathHay
75
0.11–0.14
0.01–0.03
BrowseComp
124
0.10–0.12
0.01–0.02
MCP-Bench
52
0.02–0.07
0.01–0.03
Appendix
Table 3: Bootstrap statistics across benchmarks. N denotes the number of sampled instances. Typical CI width and bootstrap standard error (Pass Std) are computed over 10,000 resamples across all evaluated models.
Task Category
Sub-category
Representative Benchmarks / Papers
Long Document QA
Single/Multi-Doc Understanding
NarrativeQA Kočiskỳ et al. [2018] , 2WikiMultihopQA Ho et al. [2020] , Loong Wang et al. [2024b]
Long Dialogue Understanding
MeetingBank Hu et al. [2023] , CharacterChat Tu et al. [2023]
Code Understanding
RepoBench Liu et al. [2023a] , CrossCodeEval Ding et al. [2023]
Structured Data Understanding
TabFact Chen et al. [2019] , WikiTableQuestions Kweon et al. [2023] , LongBench Bai et al. [2024] (L-Data)
Summarization
Global Information Aggregation
GovReport Huang et al. [2021] , Multi-News Fabbri et al. [2019] , ZeroSCROLLS Shaham et al. [2023]
Many-Shot
In-Context Learning
MIR-Bench Yan et al. [2025] , Many-Shot ICL Agarwal et al. [2024]
Appendix
Table 4: Taxonomy of Existing Long-Context Benchmarks.
Models
LongBench
HELMET
MRCR
Avg
Open-Source
GPT-OSS-120B
47.8
12.9
32.8
31.2
Qwen3-235B-A22B
58.3
40.8
40.6
46.6
Qwen3-Next
53.1
26.3
27.9
35.8
DeepSeek-V3.2
50.3
48.0
33.2
43.8
DeepSeek-R1
58.3
36.9
39.2
44.8
Appendix
Table 5 : Performance on long-context benchmarks. We report scores across three established benchmarks focusing on retrieval and single-turn comprehension. Bold marks the best and underline the second-best score in each column.
Figure 10 : Pairwise correlation between static long-context benchmarks and agentic domains in our General AgentBench. “All” denotes the average performance across the search, code, reasoning, and tool-use domains.
Figure 11 : Extended parallel scaling to K=8 and K=16 . We evaluate Gemini-2.5-Flash and DeepSeek-V3.2 on coding and tool-call domains with three commitment strategies and GPT-5 as an external verifier. The verification gap persists across all K values; majority voting growth decelerates while pairwise selection shows an accelerating trend.
Verifier
Run 1
Run 2
Run 3
Std
Self-judge (pointwise)
0.1795
0.1823
0.1705
0.0050
Self-judge (pairwise)
0.1795
0.1795
0.1672
0.0058
GPT-5 (pointwise)
0.1397
0.1471
0.1397
0.0035
Appendix
Table 6: Verification accuracy across three independent runs (Gemini-2.5-Flash, coding domain). Low standard deviations confirm that both self-judgment and external verification are stable across runs.
Benchmark
Out-of-domain tool calls (%)
SWE-Bench
∼ 5
Terminal-Bench
∼ 9
MCP-Bench
∼ 8
BrowseComp
∼ 13
Tau2-Bench
∼ 40
WebVoyager
∼ 50
Appendix
Table 7: Proportion of out-of-domain tool calls per benchmark under the unified setting (Gemini-2.5-Flash).
Tool registry size
24K
32K
36K
42K (Original)
Gemini-2.5-Flash
0.3550
0.3594
0.3572
0.3478
Appendix
Table 8: Performance of Gemini-2.5-Flash on the tool-call domain as the tool registry size varies.
Model
Input
Cached Input
Output
Gemini-2.5-Flash
0.30
0.03
2.50
Gemini-2.5-Pro
1.25
0.125
10.00
GPT-5
1.25
0.125
10.00
Claude-Haiku 4.5
1.00
0.50
5.00
Claude-Sonnet 4.5
3.00
3.75
15.00
gpt-oss-120B
0.15
0.075
0.60
Appendix
Table 9: Unit API prices (USD per 1M tokens) used in our cost estimation.
Model
Search
MathHay
SWEBench
MCPBench
Tau2Bench
TerminalBench
Total
Gemini-2.5-Flash
$193
$70.9
$12719
$122
$62.2
$826
$13993
DeepSeek-R1
$5188
$157
$3675
$492
$202
$1505
$11218
DeepSeek-V3.2
$369
$38.8
$7.80
$88.4
$54.4
$412
$970
Qwen3-235B
$1254
$41.4
$1286
$120
$55.5
$409
$3166
Qwen3-Next
$25.3
$9.52
$3.38
$21.2
$14.9
$154
$229
Total
$7028
$317
$17692
$843
$389
$3307
$29576
Appendix
Table 10 : Cost for evaluating models under parallel scaling setting (USD)
Model
Search
MathHay
SWEBench
MCPBench
Tau2Bench
TerminalBench
Total
Gemini-2.5-Flash
$5588
$369
$2024
$568
$852
$1870
$11271
DeepSeek-R1
$1267
$654
$892
$931
$1088
$782
$5614
DeepSeek-V3.2
$902
$333
$191
$79.1
$52.3
$356
$1913
Qwen3-235B
$1370
$238
$436
$716
$162
$689
$3610
Qwen3-Next
$442
$219
$392
$152
$251
$529
$1985
Total
$9568
$1814
$3935
$2445
$2405
$4225
$24392
Appendix
Table 11 : Cost for evaluating models under sequential scaling setting (USD)
Model
Search
MathHay
SWEBench
MCPBench
Tau2Bench
TerminalBench
Total
Gemini-2.5-Pro
$47.5
$18.6
$3253
$31.3
$15.2
$203
$3569
GPT-5
$87.5
$14.2
$146
$21.0
$12.4
$60.9
$342
Claude-Haiku-4.5
$304
$10.9
$312
$29.3
$14.1
$106
$776
Claude-Sonnet-4.5
$1248
$40.9
$926
$129
$52.0
$376
$2772
OpenAI-oss-120B
$1.44
$1.49
$0.52
$0.50
$1.35
$0.40
$5.70
DeepSeek-V3.2
$96.8
$10.0
$1.92
$22.0
$14.0
$102
$247
Appendix
Table 12 : Cost for evaluating models under general (default context) setting (USD)
Context
DeepSeek-R1
Qwen3-Next
Qwen3-235B
DeepSeek-V3.2
Gemini2.5-Flash
Search
64k
0.2060
0.2186
0.0510
0.0854
0.1558
80k
0.1210
0.2111
0.1256
0.2663
0.1658
96k
0.0550
0.2312
0.1457
0.2410
0.1759
112k
0.1260
0.2663
0.2010
0.2320
0.1759
128k
0.1260
0.2500
0.1809
0.2320
0.1759
Appendix
Table 13 : Detailed numerical results for sequential scaling (Figure 3 ). Accuracy of 5 models across 4 domains at increasing context budgets.
Context
Qwen3-235B
Gemini-2.5-Flash
Memory
Verifier
Ours
Memory
Verifier
Ours
Tool-call
64k
0.446
0.473
0.458
0.477
0.504
0.489
80k
0.462
0.491
0.474
0.326
0.454
0.338
96k
0.479
0.489
0.471
0.317
0.446
0.329
128k
0.449
0.488
0.441
0.317
0.387
0.329
Appendix
Table 14: Detailed numerical results for the ablation on sequential scaling strategies (Figure 6 ). Three protocols—Memory (summarization), Verifier (verifier-guided revision), and Ours (naive scaling)—are compared on two models across two domains at five context budgets.
Model
K
Pass@ K
Majority
Self-Choice
Pairwise
GPT5
DeepSeek-R1
1
0.2010
0.2010
0.2010
0.2010
0.0704
2
0.2460
0.2202
0.2111
0.2111
0.0754
3
0.2760
0.2387
0.2161
0.2161
0.0955
4
0.2970
0.2286
0.2211
0.2211
0.1156
Qwen3-Next
1
0.1859
0.1859
0.1859
0.1859
0.0352
2
0.2462
0.2167
0.1357
0.2010
0.0653
Appendix
Table 15: Verification gap – Search (Figure 7 ). Pass@ K (oracle upper bound) and four commitment strategies at K=1,2,3,4 .
Model
K
Pass@ K
Majority
Self-Choice
Pairwise
GPT5
DeepSeek-R1
1
0.1076
0.1076
0.1076
0.1076
0.1084
2
0.1690
0.1520
0.1548
0.1640
0.1626
3
0.2462
0.2312
0.2249
0.2265
0.2094
4
0.2620
0.2380
0.2483
0.2343
0.2249
Qwen3-Next
1
0.0923
0.0923
0.0923
0.0923
0.0929
2
0.1385
0.1335
0.1162
0.1239
0.1239
Appendix
Table 16: Verification gap – Code (Figure 7 ).
Model
K
Pass@ K
Majority
Self-Choice
Pairwise
GPT5
DeepSeek-R1
1
0.4667
0.4667
0.4667
0.4667
0.4267
2
0.5600
0.4932
0.4267
0.4800
0.5067
3
0.5867
0.4824
0.4933
0.4533
0.4933
4
0.6133
0.4954
0.5200
0.4400
0.5178
Qwen3-Next
1
0.4200
0.4200
0.4200
0.4200
0.4000
2
0.6000
0.5368
0.4667
0.5330
0.4733
Appendix
Table 17: Verification gap – Reason (Figure 7 ).
Model
K
Pass@ K
Majority
Self-Choice
Pairwise
GPT5
DeepSeek-R1
1
0.3968
0.3968
0.3968
0.3968
0.2200
2
0.4716
0.4611
0.4284
0.4446
0.2800
3
0.4873
0.4621
0.4297
0.4470
0.2800
4
0.5331
0.5061
0.4748
0.4912
0.3050
Qwen3-Next
1
0.5678
0.5678
0.5678
0.5678
0.1800
2
0.6418
0.5969
0.5775
0.5842
0.2800
Appendix
Table 18: Verification gap – Tool-use (Figure 7 ).
Figure 12 : Attention behavior of Qwen3-235B (full attention) under the General AgentBench setting.
Figure 13 : Attention behavior of Qwen3-Next (hybrid linear attention) under the General AgentBench setting.