cs.AIFeb 22, 2026

Evaluating Test-Time Scaling of General LLM Agents

Authors: Xiaochuan Li, Ryan Ming, Pranav Setlur, Abhijay Paladugu, Andy Tang, Hao Kang, Shuai Shao, Rong Jin, +1 more

Organizations: Language Technologies Institute, School of Computer Science, Carnegie Mellon University · Meta

Abstract

LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood. In this paper, we systematically investigate two principal test-time scaling axes of LLM agents: sequential scaling through extended interaction and parallel scaling through trajectory sampling. We first introduce a realistic benchmark that provides one unified framework for evaluating LLM agents across search, coding, reasoning, and tool-use domains, more faithfully reflecting the heterogeneity of real-world deployments. Evaluating ten leading LLM agents reveals substantial performance degradation when transitioning from domain-specific evaluations to this realistic setting. Building on this foundation, we progressively scale test-time compute along fine-grained increments to characterize the performance upper bound. We find that neither scaling axis can consistently yield meaningful gains from additional test-time compute in realistic environments, a phenomenon we attribute to two fundamental limitations: the scaling plateau that bottlenecks sequential scaling and the verification gap that undermines parallel scaling. Code is publicly available at https://github.com/cxcscmu/General-AgentBench.

Figures & tables

Appendix figures & tables22 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

    Sep 14, 2026Kaiyuan Liu, Qiuyang Mang, Bo Peng +6Large Language Model Agents

  2. Scaling Test-Time Compute for Agentic Coding

    Apr 16, 2026Joongwon Kim, Wannan Yang, Kelvin Niu +13Long-Horizon AgentsCoding Agents

  3. LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling

    May 8, 2026Tong Zheng, Haolin Liu, Chengsong Huang +10Test-Time ScalingAgentic Discovery