cs.AIOct 1, 2026

Agents Are Systems, Not Models: Rethinking Agentic Evaluation

Authors: Luis Wiedmann, Leander Girrbach, Cordelia Schmid, Zeynep Akata

Organizations: Technical University of Munich, MCML · Helmholtz Munich · Inria, École normale supérieure, CNRS, PSL Research University

Abstract

Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.

Figures & tables

Appendix figures & tables25 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

    Sep 30, 2026Michael Hardy, Ruhana Azam, Anka Reuel +2Agentic EvaluationsScaffolds

  2. Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

    May 27, 2026Yilun Yao, Xinyu Tan, Chao-Hsuan Liu +9Agent HarnessAgentic Benchmarks

  3. AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

    May 19, 2026Parsa Mazaheri, Kasra MazaheriLarge Language Model AgentsLeaderboard