TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation
Organizations: Prime Intellect
Abstract
Verifier scores now serve as both benchmark metrics and training rewards for large language model (LLM) agents, and a change in score is routinely read as a change in capability. It may instead reflect a change in the evaluation. We introduce TRACE, a protocol that turns a score change from a verdict into a testable diagnosis: it applies a targeted change to one part of an evaluation, compares paired runs, checks whether the agent's behavior changed, and rescores unchanged trajectories to test whether the scoring rule is responsible. In a controlled suite of 25 synthetic tasks, renaming tools lowers a scripted agent's score by 0.250 even though it performs exactly the same operations; restoring the original names at scoring time closes the entire gap, while the same mutation exposes a genuine behavioral failure in a second agent. On public -bench tasks with four LLM agents, an initial 30-task study finds mixed reward changes whose one clear effect does not replicate. In a larger follow-up on 88 new tasks with repeated runs per condition, renaming tools or reformatting tool outputs leaves reward unchanged to within 0.10 for seven of eight agent-change pairs, whereas tool names that deliberately mislead lower every agent's reward by 0.20-0.44, showing that the setup can detect real effects. Identical reruns flip 15-36% of task outcomes, so single-run comparisons cannot separate presentation effects from run-to-run variation. Two frontier LLM judges give consistent verdicts when a fixed trajectory is presented differently, yet disagree with each other on 57% of the same records, largely because one grades procedure rather than outcome. TRACE thus separates what a score change says about the agent from what it says about the measurement.
Figures & tables
| Probe | Implemented change | Interpretation |
|---|---|---|
| Task wording | Prepend a fixed instruction to preserve constraints. | Limited wording check; not a diverse paraphrase test. |
| Tool interface | Rename tools while preserving dispatch and arguments. | Check whether equivalent operations receive equivalent credit. |
| Rubric | Change component weights and require an evidence mention. | Measure sensitivity to a different scoring rule. |
| Added constraint | Require evidence citation in the final report. | Test an additional reporting requirement. |
| Proxy incentive | Increase visible incentives on applicable reward-hacking tasks. | Compare the proxy with simulator task quality. |
| Observation access | Return partial experiment results; expose details through log inspection. | Test behavior when evidence requires an extra access step. |
| Policy | Original | Mutated | Gap | MSR | MTR | Proxy PTD |
|---|---|---|---|---|---|---|
| Careful | 0.807 | 0.766 | 0.041 | 0.860 | 0.860 | |
| Overfit | 0.549 | 0.526 | 0.023 | 0.053 | 0.667 | 0.940 |
| Shortcut | 0.345 | 0.339 | 0.006 | 0.000 | — | 0.932 |
| Agent | Condition | Pairs | Original | Mutated | 95% task-bootstrap CI | |
|---|---|---|---|---|---|---|
| DeepSeek (internal) | Alias | 29/30 | .828 | .828 | .000 | [-.172, .172] |
| DeepSeek (internal) | Format | 29/30 | .828 | .862 | +.034 | [-.103, .172] |
| GLM (internal) | Alias | 30/30 | .833 | .767 | -.067 | [-.267, .133] |
| GLM (internal) | Format | 30/30 | .833 | .833 | .000 | [-.167, .167] |
| Laguna (internal) | Alias | 30/30 | .500 | .767 | +.267 | [.067, .433] |
| Laguna (internal) | Format | 30/30 | .500 | .633 | +.133 | [-.067, .333] |
| Claim to test | What can go wrong | TRACE evidence | Remaining limit |
|---|---|---|---|
| Equal work, equal credit | Wording or tool-surface dependence | Task and tool probes; alias-normalized rescoring removes careful’s tool gap ( ) | Synthetic tools; two presentation changes |
| Reward reflects task quality | Visible proxy exploitation | Proxy-incentive probes; shortcut proxy PTD on the 12 proxy-relevant tasks | Hand-designed proxy channels |
| Stable findings | Threshold, task-sample, or rerun sensitivity | MSR/MTR at several thresholds; replication with rerun and positive controls, analysis fixed in advance | margin; four agents |
| Evidence beyond scripts | Synthetic-only evidence | Paired studies on 30, 40, and 158 tasks with native evaluator; repeat-controlled judge audits | No judge ground truth; no training loop |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| Selection objective | Selected template | MSR | Proxy PTD |
|---|---|---|---|
| Visible proxy | proxy_exploit | 0.103 | 0.996 |
| Tool volume | careful | 0.718 | |
| Composite score | careful | 0.718 |
| Agent | Change | Tasks | Original | 95% CI | 90% CI | Label | Rerun flips | Excess flips [95% CI] | |
| DeepSeek (internal) | Alias | 87 | .843 | -.027 | [-.068, +.011] | [-.061, +.004] | equiv. | .147 | -.031 [-.077, +.012] |
| DeepSeek (internal) | Format | 86 | .841 | -.016 | [-.054, +.020] | [-.049, +.016] | equiv. | .147 | -.023 [-.065, +.020] |
| GLM (internal) | Alias | 88 | .831 | -.002 | [-.053, +.051] | [-.044, +.042] | equiv. | .157 | -.028 [-.070, +.015] |
| GLM (internal) | Format | 88 | .833 | +.030 | [-.011, +.072] | [-.008, +.064] | equiv. | .159 | -.023 [-.068, +.027] |
| Laguna (internal) | Alias | 88 | .610 | +.004 | [-.061, +.068] | [-.053, +.057] | equiv. | .356 | -.019 [-.091, +.049] |
| Laguna (internal) | Format | 88 | .610 | .000 | [-.068, +.068] | [-.057, +.057] | equiv. | .356 | +.015 [-.053, +.080] |
| Agent | Condition | Pairs | Original | Mutated | 95% CI | Switches | |
|---|---|---|---|---|---|---|---|
| DeepSeek (internal) | Alias | 39/40 | .846 | .769 | -.077 | [-.179, +.026] | 1 / 4 |
| DeepSeek (internal) | Format | 39/40 | .846 | .821 | -.026 | [-.128, +.077] | 2 / 3 |
| GLM (internal) | Alias | 40/40 | .900 | .850 | -.050 | [-.150, +.050] | 1 / 3 |
| GLM (internal) | Format | 40/40 | .900 | .850 | -.050 | [-.150, +.025] | 1 / 3 |
| Laguna (internal) | Alias | 40/40 | .650 | .600 | -.050 | [-.225, +.125] | 6 / 8 |
| Laguna (internal) | Format | 40/40 | .650 | .750 | +.100 | [-.050, +.250] | 7 / 3 |
| Judge | View | Events | Disagreement [95% CI] | Excess over repeat, pp [95% CI] |
|---|---|---|---|---|
| GPT-6.1 Sol | Repeat | 0/354 | 0.00% [—] | |
| GPT-6.1 Sol | Alias | 1/354 | 0.28% [0.00, 0.85] | +0.28 [0.00, +0.85] |
| GPT-6.1 Sol | Format | 2/354 | 0.56% [0.00, 1.71] | +0.56 [0.00, +1.71] |
| Claude Opus 5.5 | Repeat | 13/353 | 3.68% [0.86, 7.56] | |
| Claude Opus 5.5 | Alias | 10/354 | 2.82% [0.85, 5.60] | -0.85 [-3.72, +1.15] |
| Claude Opus 5.5 | Format | 12/354 | 3.39% [1.15, 5.93] | -0.28 [-2.82, +2.23] |
| Agent | Change | Tasks | Original | 95% CI | 90% CI | Label | Rerun flips | Excess flips [95% CI] | |
|---|---|---|---|---|---|---|---|---|---|
| DeepSeek (internal) | Alias | 108 | .759 | -.035 | [-.088, +.017] | [-.080, +.009] | equiv. | .185 | +.009 [-.037, +.059] |
| DeepSeek (internal) | Format | 108 | .756 | -.023 | [-.074, +.028] | [-.068, +.020] | equiv. | .187 | +.019 [-.035, +.074] |
| GLM (internal) | Alias | 108 | .994 | +.003 | [-.006, +.012] | [-.006, +.012] | equiv. | .012 | -.003 [-.012, +.006] |
| GLM (internal) | Format | 108 | .994 | -.015 | [-.037, +.003] | [-.034, .000] | equiv. | .012 | +.015 [-.003, +.037] |
| Laguna (internal) | Alias | 108 | .404 | +.006 | [-.052, +.065] | [-.043, +.056] | equiv. | .309 | -.019 [-.068, +.031] |
| Laguna (internal) | Format | 108 | .404 | -.056 | [-.114, +.003] | [-.105, -.006] | inconcl. | .309 | -.037 [-.086, +.019] |