Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation
Organizations: Khoury College of Computer Sciences Northeastern University San Jose, CA, USA · School of Engineering and Applied Science The George Washington University Bellevue, WA, USA · Khoury College of Computer Sciences Northeastern University Boston, MA, USA · Department of Computer Science University of California, Berkeley Berkeley, CA, USA
Abstract
Agentic leaderboards increasingly evaluate systems on public benchmarks whose task statements and solution-bearing artifacts can remain accessible. We propose a measurement-first audit framework that distinguishes contamination claims according to the evidence required to support them. It separates three channels that require different evidence: training-time exposure, evaluation-time retrieval, and pipeline/scaffold leakage. Each channel is coded as open, partial, closed, or unknown under a fail-closed rule. Across nine Holistic Agent Leaderboard (HAL) configurations, none of the 27 channel assessments was coded closed, but incidents were confirmed in four configurations. We then apply the behavioral component of the framework to a reported file-localization gap on SWE-bench Verified, using an outcome-blind, same-repository matched-control design with symmetric prompt-leakage screening, paired and repository-aware uncertainty analyses, and scorer validation, evaluated on GPT-4.1 and DeepSeek-V4-Flash. Among the 100 pairs retained after symmetric screening and the pair-integrity exclusion, GPT-4.1 showed a -point pair-weighted Top-3 benchmark-associated gap, but the 95% intervals from both the prespecified paired-bootstrap procedure and the post-hoc repository-balanced analysis included zero, leaving the benchmark-associated gap inconclusive. The reproduction scorer did not pass its validation gate: against consensus human labels, sufficient scorer sensitivity could not be established for either model, and both DeepSeek-V4-Flash firings on correct-gold comparisons were false positives. Without provenance evidence, appropriate controls, symmetric leakage screening, and validated scorers, stronger contamination claims are not warranted. The results do not establish training-data membership, contamination prevalence, or benchmark-induced score inflation.
Figures & tables
| Susceptibility | Realized incident | ||||
|---|---|---|---|---|---|
| Benchmark | T | R | L | Confirmed | Revision |
| Online Mind2Web (BrowserUse / SeeAct) | P [Y/N/N/?] | O | ? | Not established | Unknown |
| AssistantBench (Browser agent) | P [Y/Y/Y/?] | O | O | L | Current |
| GAIA (Open Deep Research) | P [Y/Y/Y/?] | O | ? | Not established | Unknown |
| CORE-Bench Hard (CORE-Agent / Generalist) | P [Y/Y/Y/?] | O | P | Not established | Unknown |
| ScienceAgentBench (SAB Self-Debug) | P [Y/Y/Y/?] | P | P | None | — |
| Model | Metric | Pair-weighted gap [95% CI] | Repository-balanced gap [95% CI] |
|---|---|---|---|
| DeepSeek-V4-Flash | Top-3 (primary) | ||
| Top-1 | |||
| MRR@3 | |||
| Path F1 | |||
| GPT-4.1 | Top-3 (primary) | ||
| Top-1 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage | Change | Pairs retained |
|---|---|---|
| Initial matched sample | — | 199 |
| Pair-level leakage exclusion | 181 | |
| Cross-arm underlying-issue collision | 180 | |
| Persistent truncation | 173 | |
| No eligible strict signature in benchmark arm | 84 | |
| No eligible strict signature in control arm | 42 |
| Model | Human-positive responses | Detected by strict scorer | Weighted sensitivity | Exact 95% CI, unweighted sensitivity | Precision |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash | 2/41 | 0/2 | 0.00 | 0/2 | |
| GPT-4.1 | 3/38 | 1/3 | 0.19 | 1/1 |
| Model | Sampling stratum | Available arms | Reviewed arms | Inclusion probability | Weight |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash | Exact hit | 2 | 2 | 1.000 | 1.00 |
| Fuzzy-only | 6 | 6 | 1.000 | 1.00 | |
| High overlap | 15 | 15 | 1.000 | 1.00 | |
| Random tail | 61 | 20 | 0.328 | 3.05 | |
| GPT-4.1 | Exact hit | 1 | 1 | 1.000 | 1.00 |
| Fuzzy-only | 4 | 4 | 1.000 | 1.00 |
| Model | N pairs | Pair-weighted gap [95% CI] | Repository-balanced gap [95% CI] | Hierarchical 95% CI |
|---|---|---|---|---|
| DeepSeek-V4-Flash | 64 | |||
| GPT-4.1 | 64 |
| Agreement unit | Agreements / total | Agreement |
|---|---|---|
| Field-level judgments | 90/117 | 76.9% |
| Derived T labels | 9/9 | 100.0% |
| Derived R labels | 6/9 | 66.7% |
| Derived L labels | 6/9 | 66.7% |
| All derived T/R/L labels | 21/27 | 77.8% |
| Incident channel/status | 4/9 | 44.4% |
| Hypothesis | Disposition | Evidence |
|---|---|---|
| H1: Benchmark tasks exhibit higher file-path accuracy. | Inconclusive; point estimates were positive, but the primary 95% intervals included zero. | All gaps were positive; no Top-3 primary interval excluded zero (Section 4.3 ). |
| H2: Benchmark tasks exhibit higher patch-specific line recovery. | Unresolved: measurement validity insufficient. | The scorer did not pass the overall validation gate: sufficient sensitivity could not be established for either model, and both DeepSeek-V4-Flash firings on correct-gold comparisons were human-negative. The post-hoc strict-score contrasts therefore remain exploratory. |
| H3: Reproduction is rarer than localization and, conditional on adequate scorer sensitivity, has an equal or larger benchmark-control gap. | Unresolved: required sensitivity condition unmet. | The sensitivity condition could not be established, so neither the reproduction-frequency comparison nor the conditional benchmark-control gap comparison could be adjudicated. |