Benchmark Auditing

Momentum

17 papers in the last four weeks, up 325% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 60

All topics
CardsList
  1. Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation

    Oct 5, 2026Dishu Yang, Qi Su, Hongbo Qin +1Benchmark AuditingAI Agent Auditing

  2. Protocol-Sensitive Evaluation of Log Anomaly Detection: Component Costs and Target-Access Sensitivity on HDFS and BGL

    Oct 5, 2026Hang Xiao, Janet Sung, Zhaoyi Li +2Benchmark AuditingAnomaly Detection

  3. SALUS: Automated Auditing of NL-to-SQL Benchmarks through Weak Supervision of Multi-Agent Output

    Oct 4, 2026Shiyuan Zhou, Ashwin Gerard Colaco, Sainyam Galhotra +1Benchmark AuditingText-to-SQL

  4. Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves

    Sep 30, 2026Sohail, Sarkar, Shakuntala BaichooBenchmark AuditingTest-Time Scaling

  5. Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

    Sep 29, 2026Rohith Reddy Bellibatlu, Zichong Wang, Wenbin ZhangTool-Use EvaluationAutomated Software Testing

  6. The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection

    Sep 27, 2026Farhan Shahriyar Hossain, Taufikur Rahman Fuad, Md Abrar Jahin +1Few-Shot Anomaly DetectionBenchmark Auditing

  7. Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark

    Sep 23, 2026Makar Ulesov, Vladislav Smirnov, Omar Ibrahim +1Benchmark AuditingLLM Auditing

  8. Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL

    Sep 22, 2026Jiamiao Liu, Dewen Qiao, Yu Zhang +1Probability CalibrationBenchmark Auditing

  9. ReSTI: A Source-Grounded Audit and Repair of STI-Bench

    Sep 21, 2026Pengzhan Sun, Ramanathan Rajaraman, Shiu-Hong Kao +2Spatial Reasoning BenchmarksSpatiotemporal Reasoning

  10. ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift

    Sep 20, 2026Kemal Kirtac, Carsten MapleDistribution Shift RobustnessBenchmark Auditing

  11. Auditing Bayesian Graph Alignment: Diagnostic Comparisons and Reference Failure

    Sep 19, 2026Melika Gorgi, Kourosh MirsohiMarkov Chain Monte CarloBenchmark Auditing

  12. DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?

    Sep 14, 2026Hongye Yang, Zhihao Xie, Shengjun Xiong +1Benchmark AuditingText-to-CAD Generation

  13. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

    Sep 10, 2026Koutian Wu, Junjie Zhou, Ergan Shang +6Benchmark AuditingBenchmark Design

  14. Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

    Aug 31, 2026Ahmed El Kady, Aravind Narayanan, Rehana Noorani +2LLM EvaluationEnergy-Efficient ML

  15. Auditing MCQA Benchmarks through Probability Landscapes

    Aug 31, 2026Minsoo Song, Chanjun ParkBenchmark AuditingLLM Auditing

  16. CatchBench: When Can an Agent Failure Be Caught?

    Aug 24, 2026Yue Zhao, Mengyuan Li, Ruolin Li +5Agent Failure AnalysisBenchmark Auditing

  17. When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

    Aug 8, 2026Ibne Farabi Shihab, Sanjeda Akter, Anuj SharmaStatistical Power AnalysisBenchmark Auditing

  18. On the missing benchmarks layer and a potential solution

    Aug 4, 2026Francis F Daniel, Mauro Ibañez, Francis Perelman +1LLM EvaluationBenchmark Auditing

  19. Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch

    Aug 3, 2026Shuyang Xie, Shuxiao Xie, Feng Zhu +2Automated Software TestingBenchmark Auditing

  20. Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

    Jul 30, 2026Philipp D. Siedler, Jordan SassoonLLM EvaluationBenchmark Auditing

  21. PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

    Jul 30, 2026Manyi Wang, Junjielong Xu, Pinjia HeBenchmark AuditingSoftware Engineering Benchmarks

  22. How Benchmarks Mis-Score Computer-Use Agents

    Jul 30, 2026Zihan Dong, Zhiyuan Ma, Zekun Wang +5Agent Failure AnalysisComputer-Use Agent Benchmarks

  23. Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks

    Jul 29, 2026Jeff Mohl, Nelson Gardner-Challis, Magda Dubois +6Benchmark AuditingAI Agent Benchmarks

  24. Visual Credit Audit for Multimodal Spatial Reasoning

    Jul 29, 2026Feixiang Liu, Qiang Qiu, Lanbo Sun +3VLM EvaluationSpatial Reasoning Benchmarks

  25. Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation

    Jul 27, 2026Jingkun Luo, Da-Tian PengData ProvenanceBenchmark Auditing

  26. Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

    Jul 24, 2026Jiaqi Shao, Hanck Chen, Wei Zhang +2Computer-Use Agent BenchmarksReward Hacking

  27. Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

    Jul 1, 2026Zhi Chen, Zhensu Sun, Yuling Shi +2Benchmark AuditingBenchmark Design

  28. Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

    Jul 1, 2026Yanhang Li, Zhichao Fan, Zexin ZhuangAI AssuranceBenchmark Auditing

  29. Auditing Generalization in AI-Generated Video Detection: A Six-Control Protocol and the VidAudit Toolkit

    Jun 30, 2026Mert Onur Cakiroglu, Zhihe Lu, Mehmet Dalkilic +1Benchmark AuditingOOD Generalization

  30. Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks

    Jun 28, 2026Lining Hu, Ting Liu, Yuzhuo FuPairwise ComparisonAlgorithmic Auditing

  31. The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

    Jun 2, 2026Wojciech Zarzecki, Jan Dubiński, Sebastian CygertBenchmark AuditingLLM Auditing

  32. Auditing LLM Benchmarks with Item Response Theory

    May 28, 2026Sander Land, Daniel M. BikelBenchmark AuditingLLM Auditing

  33. FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

    May 27, 2026Nishal Thomas, Noel ThomasMathematical Reasoning BenchmarksBenchmark Auditing

  34. Automated Benchmark Auditing for AI Agents and Large Language Models

    May 25, 2026Junlin Wang, Federico Bianchi, Shang Zhu +4LLM EvaluationBenchmark Auditing

  35. Deployment-complete benchmarking

    May 25, 2026El Mustapha Mansouri, Keigo AraiBenchmark AuditingBenchmark Construction

  36. Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

    May 25, 2026Zexin Zhuang, Yanhang Li, Zhichao FanLLM QuantizationStatistical Power Analysis

  37. What Twelve LLM Agent Benchmark Papers Disclose About Themselves: A Pilot Audit and an Open Scoring Schema

    May 20, 2026Mahdi Naser Moghadasi, Faezeh GhaderiBenchmark AuditingLLM Auditing

  38. AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

    May 19, 2026Parsa Mazaheri, Kasra MazaheriAgent Failure AnalysisComputer-Use Agent Benchmarks

  39. Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks

    May 14, 2026Karthik Raghu Iyer, Yazdan Jamshidi, Nicholas Bray +1LLM Safety BenchmarksBenchmark Auditing

  40. Do Composed Image Retrieval Benchmarks Require Multimodal Composition?

    May 14, 2026Matteo Attimonelli, Alessandro De Bellis, Aryo Pradipta Gema +8Benchmark AuditingImage Retrieval

  41. BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

    Apr 27, 2026Xinming Tu, Tianze Wang, Yingzhou +4Benchmark AuditingAI Agent Benchmarks

  42. AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors

    Feb 26, 2026Abhay Sheshadri, Aidan Ewart, Elias Kempf +8Benchmark AuditingAI Agent Evaluation

  43. Who Benchmarks the Benchmarks? Towards Comprehensive Evaluation of Commonsense Reasoning Benchmarks

    Apr 10, 2025Pavel Chizhov, Anton Changalidis, Vishnu Prasad Vijaya Kumar +4Benchmark AuditingReasoning Evaluation