Agents Are Systems, Not Models: Rethinking Agentic Evaluation
Organizations: Technical University of Munich, MCML · Helmholtz Munich · Inria, École normale supérieure, CNRS, PSL Research University
Abstract
Agent evaluations increasingly go beyond a single success rate, reporting metrics such as cost, consistency, and robustness. Yet they typically treat the agent itself as fixed. In practice, an agent is a configurable system: users decide what to tell it, how long to let it run, and which model to use, and each of these choices can change how well and how consistently it performs. We study these choices on a new benchmark of four scientific tasks, where a coding agent must find and correctly operate a published specialist model. We investigate five parts of the agent's configuration: task information, reasoning, self-verification, time budget, and backbone model. We find substantial run-to-run variability, with approximately 54% of the outcome variance coming from repeating the same configuration rather than changing it. Across configurations, the information provided to the agent has the largest effect, exceeding both time budget and model size, while also reducing cost and improving calibration. Configuration choices also interact: additional time helps only when the agent has sufficient information or a capable enough model to use it. Finally, a trajectory-based taxonomy of agent behavior reveals that prompting an agent to verify its answer has little effect on its verification behavior, whereas providing a dedicated verification tool changes that behavior substantially. These results suggest that agents should be evaluated as configurable systems themselves, and that some desired behaviors are more effectively implemented in the system than requested through prompting. We release the benchmark and more than 18,000 agent trajectories.
Figures & tables
| redshift- estimation | mmlu- astronomy | promoter- prediction | rna-folding | |
|---|---|---|---|---|
| Domain | Astrophysics | Astrophysics | Genomics | Genomics |
| Size | 20 galaxy images | 152 questions | 613 DNA sequences | 300 RNA sequences |
| Metric | Accuracy | MCC | ||
| Specialist | AstroCLIP | AstroSage-8B | DNABERT-2 | RiNALMo |
| Challenge | Pre-processing | Declining | Fine-tuning | Post-processing |
| Regime | gap_positive | gap_negative | gap_positive | gap_positive |
| redshift- | mmlu- | promoter- | rna- | |
| estimation | astronomy | prediction | folding | |
| Runs & Completion | ||||
| # Runs | 2,160 | 2,160 | 2,160 | 2,160 |
| Completion Rate | 76.3% 1.8% | 95.3% 0.9% | 89.3% 1.3% | 61.7% 2.0% |
| % Cells that (always / mixed / never) complete | 54.6 / 40.3 / 5.1 | 81.9 / 18.1 / 0.0 | 63.7 / 36.1 / 0.2 | 21.3 / 72.0 / 6.7 |
| Anchors | ||||
| promoter-prediction | redshift-estimation | rna-folding | |||||
|---|---|---|---|---|---|---|---|
| Axis | Rank | Effect | Rank | Effect | Rank | Effect | Mean Rank |
| Information | 1 | 1.925 | 1 | 1.452 | 1 | 1.496 | 1.00 |
| Model | 2 | 0.704 | 4 | 0.170 | 2 | 1.253 | 2.67 |
| Reasoning | 3 | 0.274 | 3 | 0.172 | 3 | 0.302 | 3.00 |
| Budget | 4 | 0.142 | 2 | 0.201 | 4 | 0.255 | 3.33 |
| Verification | 5 | 0.131 | 5 | 0.130 | 5 | 0.077 | 5.00 |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Grid | Cells per task | Runs per task | Trajectories |
|---|---|---|---|
| Main experiments (3 Qwen3.5 models) | 432 | 2,160 | 8,640 |
| Step-3.7-Flash (no act-only ) | 96 | 480 | 1,920 |
| Claude Sonnet 5 | 144 | 720 | 2,880 |
| Oracle, Qwen-122B | 144 | 720 | 2,880 |
| Oracle, Step-3.7-Flash | 96 | 480 | 1,920 |
| Total | 18,240 |
| run_bash | read_file | write_file | web_fetch | finish | Other/malformed | Total | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Task | Calls | Err.% | Calls | Err.% | Calls | Err.% | Calls | Err.% | Calls | Err.% | Calls | Err.% | Calls | Err.% |
| redshift-estimation | 28.51 | 11.4 | 4.24 | 5.7 | 3.56 | 3.9 | 0.79 | 0.0 | 0.76 | 0.1 | 0.07 | 100.0 | 37.93 | 9.7 |
| mmlu-astronomy | 17.84 | 12.7 | 2.92 | 1.3 | 2.36 | 1.4 | 0.07 | 3.8 | 0.95 | 0.0 | 0.05 | 100.0 | 24.20 | 9.9 |
| promoter-prediction | 16.31 | 16.2 | 3.09 | 0.7 | 3.73 | 3.7 | 0.04 | 4.3 | 0.89 | 0.1 | 0.04 | 100.0 | 24.10 | 11.8 |
| rna-folding | 33.21 | 14.2 | 3.73 | 4.0 | 6.14 | 3.4 | 1.62 | 0.7 | 0.62 | 0.3 | 0.18 | 100.0 | 45.49 | 11.6 |
| Task | Fetches/run | Ref. correct % | Ref. decoy % | Off-list % | HTTP OK % | Top off-list hosts |
|---|---|---|---|---|---|---|
| redshift-estimation | 0.79 | 88.9 | 0.1 | 11.0 | 84.4 | colab.research.google.com , lite.duckduckgo.com |
| mmlu-astronomy | 0.07 | 0.0 | 73.0 | 27.0 | 87.4 | lite.duckduckgo.com , en.wikipedia.org |
| promoter-prediction | 0.04 | 84.8 | 0.0 | 15.2 | 85.9 | huggingface.co , raw.githubusercontent.com |
| rna-folding | 1.62 | 95.9 | 0.0 | 4.1 | 86.7 | api.github.com , lite.duckduckgo.com |
| Task | Specialist | mentioned | read_docs | executed | researched_online |
|---|---|---|---|---|---|
| redshift-estimation | astroclip | 100.0% | 99.9% | 99.3% | 28.3% |
| astrosage | 27.3% | 27.1% | 0.2% | 0.1% | |
| dnabert-2 | 0.0% | 0.0% | 0.0% | 0.0% | |
| rinalmo | 0.0% | 0.0% | 0.0% | 0.0% | |
| mmlu-astronomy | astroclip | 30.5% | 30.2% | 0.3% | 0.0% |
| astrosage | 100.0% | 98.6% | 66.1% | 5.0% |
| Task | Hurdle failure rate (95% CI) | ICC, all runs | ICC, clears hurdle | Mean per-cell SD | 95% CI (5 runs) | |
|---|---|---|---|---|---|---|
| redshift-estimation | 53.8% [51.4, 56.2] | 0.727 | 33.3% | 36.2% | 85.1% | 91.9% |
| mmlu-astronomy | 4.7% [3.9, 5.7] | 0.905 | 29.4% | 6.9% | 17.6% | 20.8% |
| promoter-prediction | 0.5% [0.2, 0.9] | 0.986 | 47.9% | 49.7% | 10.7% | 12.9% |
| rna-folding | 8.7% [7.3, 10.3] | 0.825 | 45.7% | 35.2% | 12.8% | 13.0% |
| Pooled (3 gap_positive tasks, task means removed) | 20.6% [19.5, 21.7] | 0.858 | 39.4% | 46.1% | — | — |
| Info/Budget | Short | Medium | Long |
|---|---|---|---|
| None | |||
| 89.1% | 92.0% | 92.2% | |
| Identity | |||
| 83.8% | 84.4% | 84.4% | |
| Interface | |||
| 76.1% | 71.9% | 65.9% |
| redshift- | mmlu- | promoter- | rna- | |
|---|---|---|---|---|
| estimation | astronomy | prediction | folding | |
| Mean Runtime (s) | 344.2 | 188.3 | 278.0 | 476.0 |
| Mean Cost (USD) | 0.277 | 0.107 | 0.115 | 0.424 |
| Mean Steps / Tool Calls / Errors | 33 / 38 / 4 | 21 / 24 / 2 | 20 / 24 / 3 | 41 / 45 / 5 |
| Qwen-35B-A3B | Qwen-122B-A10B | Qwen-397B-A17B | Sonnet 5 | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Task | Tokens | Cost | Tokens | Cost | Tokens | Cost | Tokens | Calculated | Real | Savings |
| redshift | 659.5 / 10.3 | $186 | 397.2 / 7.7 | $183 | 340.1 / 7.0 | $229 | 147.8 / 4.1 | $337 | $124 | 63.3% |
| mmlu | 222.3 / 3.7 | $63 | 142.0 / 3.2 | $67 | 150.7 / 3.1 | $102 | 65.4 / 2.8 | $159 | $66 | 58.6% |
| promoter | 306.0 / 7.1 | $91 | 143.0 / 4.9 | $73 | 117.4 / 4.1 | $85 | 88.5 / 3.9 | $216 | $84 | 60.9% |
| rna | 978.5 / 15.3 | $275 | 579.3 / 13.0 | $273 | 530.3 / 13.5 | $367 | 140.8 / 4.6 | $328 | $110 | 66.4% |
| Outcome | Cal. Err. | Cost (USD) | Wall-clock (s) | Steps / TC / Err. | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Model | Task(s) | Baseline | Oracle | Baseline | Oracle | Baseline | Oracle | Baseline | Oracle | Baseline | Oracle |
| Qwen-122B | mmlu-astronomy † | 0.707 | 0.725 | 0.152 | 0.002 | 0.093 | 0.125 | 165 | 220 | 19 / 22 / 2 | 21 / 24 / 2 |
| Pooled gap_positive | 0.822 | 0.923 | 0.093 | 0.002 | 0.245 | 0.358 | 346 | 453 | 28 / 32 / 4 | 34 / 39 / 5 | |
| Step-3.7-Flash | mmlu-astronomy † | 0.677 | 0.721 | 0.285 | 0.065 | 0.043 | 0.091 | 164 | 309 | 18 / 20 / 2 | 26 / 28 / 3 |
| Pooled gap_positive | 0.865 | 0.889 | 0.073 | 0.026 | 0.126 | 0.177 | 369 | 455 | 29 / 31 / 4 | 35 / 37 / 5 | |
| promoter | redshift | rna | Mean Rank | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Axis | Baseline | Oracle | Baseline | Oracle | Baseline | Oracle | Baseline | Oracle |
| Qwen-122B | Information | 1 | 1 | 1 | 1 | 1 | 1 | 1.00 | 1.00 |
| Reasoning | 2 | 2 | 3 | 2 | 2 | 2 | 2.33 | 2.00 | |
| Verification | 3 | 4 | 4 | 4 | 3 | 3 | 3.33 | 3.67 | |
| Budget | 4 | 3 | 2 | 3 | 4 | 4 | 3.33 | 3.33 | |
| Step-3.7-Flash | Information | 1 | 2 | 1 | 1 | 1 | 1 | 1.00 | 1.33 |
| promoter | redshift | rna | ||||||
|---|---|---|---|---|---|---|---|---|
| Model | Axis | Rank | Effect | Rank | Effect | Rank | Effect | Mean Rank |
| Qwen-122B | Information | 1 | 2.172 | 1 | 2.027 | 1 | 1.410 | 1.00 |
| Verification | 2 | 0.256 | 3 | 0.195 | 3 | 0.193 | 2.67 | |
| Budget | 4 | 0.086 | 2 | 0.313 | 2 | 0.262 | 2.67 | |
| Reasoning | 3 | 0.116 | 4 | 0.028 | 4 | 0.041 | 3.67 | |
| Step-3.7-Flash | Information | 1 | 0.848 | 1 | 0.951 | 1 | 1.412 | 1.00 |
| Model | Task(s) | Outcome | Cal. Err. | Cost (USD) | Wall-clock (s) | Steps / TC / Err. |
|---|---|---|---|---|---|---|
| Qwen-122B | mmlu-astronomy † | 0.702 | 0.153 | 0.089 | 164 | 17.8 / 21.6 / 2.1 |
| Pooled gap_positive | 0.802 | 0.105 | 0.228 | 347 | 25.7 / 31.9 / 3.7 | |
| Step-3.7-Flash | mmlu-astronomy † | 0.677 | 0.285 | 0.043 | 164 | 17.8 / 19.9 / 2.1 |
| Pooled gap_positive | 0.865 | 0.073 | 0.126 | 369 | 28.9 / 30.8 / 4.1 |
| promoter | redshift | rna | ||||||
|---|---|---|---|---|---|---|---|---|
| Model | Axis | Rank | Effect | Rank | Effect | Rank | Effect | Mean Rank |
| Qwen-397B | Information | 1 | 1.884 | 1 | 1.126 | 1 | 1.051 | 1.00 |
| Reasoning | 3 | 0.467 | 2 | 0.238 | 2 | 0.233 | 2.33 | |
| Budget | 2 | 0.497 | 4 | 0.106 | 4 | 0.070 | 3.33 | |
| Verification | 4 | 0.235 | 3 | 0.207 | 3 | 0.178 | 3.33 | |
| Claude Sonnet 5 | Budget | 1 | 1.589 | 1 | 1.878 | 1 | 1.796 | 1.00 |
| Model | Task(s) | Outcome | Cal. Err. | Cost (USD) | Wall-clock (s) | Steps / TC / Err. |
|---|---|---|---|---|---|---|
| Qwen-397B | mmlu-astronomy † | 0.715 | 0.131 | 0.141 | 159 | 19.9 / 25.0 / 2.8 |
| Pooled gap_positive | 0.870 | 0.092 | 0.315 | 342 | 24.7 / 31.4 / 2.6 | |
| Claude Sonnet 5 | mmlu-astronomy † | 0.679 | 0.043 | 0.220 | 210 | 11.5 / 11.6 / 0.3 |
| Pooled gap_positive | 0.828 | 0.016 | 0.407 | 244 | 13.5 / 14.2 / 0.5 |
| promoter-prediction | redshift-estimation | rna-folding | ||||
|---|---|---|---|---|---|---|
| Axis | P(1) | P(1) | P(1) | |||
| Information | 1.00 | 0.413 [0.378, 0.445] | 1.00 | 0.296 [0.281, 0.323] | 1.00 | 0.260 [0.231, 0.289] |
| Model | 0.00 | 0.080 [0.061, 0.100] | 0.00 | 0.007 [0.002, 0.016] | 0.00 | 0.145 [0.117, 0.177] |
| Reasoning | 0.00 | 0.011 [0.004, 0.019] | 0.00 | 0.008 [0.002, 0.017] | 0.00 | 0.008 [0.003, 0.018] |
| Budget | 0.00 | 0.009 [0.004, 0.017] | 0.00 | 0.000 [0.000, 0.004] | 0.00 | 0.004 [0.000, 0.012] |
| Verification | 0.00 | 0.002 [0.000, 0.008] | 0.00 | 0.003 [0.001, 0.012] | 0.00 | 0.000 [0.000, 0.006] |
| All completed runs | Matched configurations ( ) | |||
|---|---|---|---|---|
| Level | Mean | Median | Mean | Median |
| Model = Qwen-35B | [ , ] | [ , ] | [ , ] | [ , ] |
| Model = Qwen-122B | [ , ] | [ , ] | [ , ] | [ , ] |
| Model = Qwen-397B | [ , ] | [ , ] | [ , ] | [ , ] |
| Information = none | [ , ] | [ , ] | [ , ] | [ , ] |
| Information = identity | [ , ] | [ , ] | [ , ] | [ , ] |
| Model | LLM s / step | Steps / min | Tokens / s | CV | LLM share | Tool share |
|---|---|---|---|---|---|---|
| Qwen-35B | 2.4 | 7.2 | 127 | 0.26 | 30% | 70% |
| Qwen-122B | 4.1 | 6.0 | 83 | 0.16 | 41% | 59% |
| Qwen-397B | 4.5 | 6.1 | 70 | 0.24 | 45% | 55% |
| Step-3.7-Flash | 4.3 | 6.0 | 94 | 0.18 | 46% | 54% |
| Claude Sonnet 5 | 4.6 | 3.6 | 84 | 0.16 | 27% | 73% |
| Model | Task | 0 | 1 | 2 | 3–4 | 5 | Mean calls | Flagged |
|---|---|---|---|---|---|---|---|---|
| Qwen-122B | redshift-estimation | 19.7% | 38.9% | 8.3% | 11.7% | 21.4% | 4.6 | 9.7% |
| mmlu-astronomy | 8.1% | 68.3% | 8.3% | 7.5% | 7.8% | 1.7 | 0.1% | |
| promoter-prediction | 3.8% | 44.9% | 11.0% | 18.9% | 21.5% | 3.0 | 2.5% | |
| rna-folding | 27.8% | 44.6% | 9.4% | 9.2% | 9.0% | 1.7 | 2.8% | |
| Step-3.7-Flash | redshift-estimation | 44.6% | 25.8% | 4.8% | 5.6% | 19.2% | 4.0 | 14.6% |
| mmlu-astronomy | 34.6% | 25.2% | 12.3% | 14.6% | 13.3% | 2.0 | 1.5% |
| Axis | Level | Completion | Mean accuracy |
|---|---|---|---|
| Information | none | 92.8% | 0.699 |
| identity | 93.9% | 0.712 | |
| interface | 98.7% | 0.710 | |
| protocol | 95.7% | 0.617 | |
| Model | Qwen-35B | 89.6% | 0.626 |
| Qwen-122B | 97.5% | 0.707 |
| Hurdle failure rate | ICC, hurdle-clearing | |||||||
|---|---|---|---|---|---|---|---|---|
| Task | ||||||||
| redshift-estimation | 48.7% | 53.4% | 53.8% | 54.1% | 35.8% | 36.2% | 36.2% | 53.3% |
| mmlu-astronomy | 4.2% | 4.3% | 4.7% | 5.5% | 6.9% | 7.1% | 6.9% | 7.2% |
| promoter-prediction | 0.2% | 0.3% | 0.5% | 0.6% | 48.7% | 48.7% | 49.7% | 50.0% |
| rna-folding | 0.2% | 7.5% | 8.7% | 9.5% | 45.8% | 35.7% | 35.2% | 34.2% |
| Pooled | 16.4% | 20.0% | 20.6% | 20.9% | 47.4% | 44.8% | 46.1% | 46.4% |
| Error Category | Execution Quality | Model Usage | Planning & Exploration | Result | Verification Behavior | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Axis | Rank | V | Rank | V | Rank | V | Rank | V | Rank | V | Rank | V |
| Information | 1 | 0.233 | 1 | 0.264 | 1 | 0.234 | 1 | 0.271 | 1 | 0.191 | 1 | 0.112 |
| Reasoning | 5 | 0.041 | 4 | 0.044 | 4 | 0.050 | 4 | 0.038 | 4 | 0.053 | 4 | 0.030 |
| Verification | 4 | 0.048 | 5 | 0.031 | 5 | 0.020 | 5 | 0.012 | 5 | 0.031 | 5 | 0.023 |
| Budget | 3 | 0.120 | 3 | 0.070 | 2 | 0.108 | 3 | 0.039 | 2 | 0.152 | 2 | 0.108 |
| Model | 2 | 0.157 | 2 | 0.138 | 3 | 0.097 | 2 | 0.084 | 3 | 0.095 | 3 | 0.067 |