cs.AIOct 5, 2026

Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation

Authors: Dishu Yang, Qi Su, Hongbo Qin, Hansong Zhang

Organizations: Khoury College of Computer Sciences Northeastern University San Jose, CA, USA · School of Engineering and Applied Science The George Washington University Bellevue, WA, USA · Khoury College of Computer Sciences Northeastern University Boston, MA, USA · Department of Computer Science University of California, Berkeley Berkeley, CA, USA

Abstract

Agentic leaderboards increasingly evaluate systems on public benchmarks whose task statements and solution-bearing artifacts can remain accessible. We propose a measurement-first audit framework that distinguishes contamination claims according to the evidence required to support them. It separates three channels that require different evidence: training-time exposure, evaluation-time retrieval, and pipeline/scaffold leakage. Each channel is coded as open, partial, closed, or unknown under a fail-closed rule. Across nine Holistic Agent Leaderboard (HAL) configurations, none of the 27 channel assessments was coded closed, but incidents were confirmed in four configurations. We then apply the behavioral component of the framework to a reported file-localization gap on SWE-bench Verified, using an outcome-blind, same-repository matched-control design with symmetric prompt-leakage screening, paired and repository-aware uncertainty analyses, and scorer validation, evaluated on GPT-4.1 and DeepSeek-V4-Flash. Among the 100 pairs retained after symmetric screening and the pair-integrity exclusion, GPT-4.1 showed a +10.0+10.0-point pair-weighted Top-3 benchmark-associated gap, but the 95% intervals from both the prespecified paired-bootstrap procedure and the post-hoc repository-balanced analysis included zero, leaving the benchmark-associated gap inconclusive. The reproduction scorer did not pass its validation gate: against consensus human labels, sufficient scorer sensitivity could not be established for either model, and both DeepSeek-V4-Flash firings on correct-gold comparisons were false positives. Without provenance evidence, appropriate controls, symmetric leakage screening, and validated scorers, stronger contamination claims are not warranted. The results do not establish training-data membership, contamination prevalence, or benchmark-induced score inflation.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

    Aug 7, 2026Sahil Pardasani, Madhusudan SinghReal-Tool BenchmarksDeepseek

  2. Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

    Jul 24, 2026Jiaqi Shao, Hanck Chen, Wei Zhang +2Agentic BenchmarksExploitation