Benchmark Auditing

Momentum

17 papers in the last four weeks, up 325% on the four weeks before. 0.2% of all new papers.

Jul 13Week of Sep 28

Latest papers 61

All topics
CardsList
  1. Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validation

    Oct 5, 2026Dishu Yang, Qi Su, Hongbo Qin +1Benchmark AuditingAI Agent Auditing

  2. Protocol-Sensitive Evaluation of Log Anomaly Detection: Component Costs and Target-Access Sensitivity on HDFS and BGL

    Oct 5, 2026Hang Xiao, Janet Sung, Zhaoyi Li +2Benchmark AuditingAnomaly Detection

  3. SALUS: Automated Auditing of NL-to-SQL Benchmarks through Weak Supervision of Multi-Agent Output

    Oct 4, 2026Shiyuan Zhou, Ashwin Gerard Colaco, Sainyam Galhotra +1Benchmark AuditingText-to-SQL

  4. Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves

    Sep 30, 2026Sohail, Sarkar, Shakuntala BaichooBenchmark AuditingTest-Time Scaling

  5. Do Agent Benchmarks Do What They Say? An Executable-Contract Audit of Tool-Using Agent Environments

    Sep 29, 2026Rohith Reddy Bellibatlu, Zichong Wang, Wenbin ZhangTool-Use EvaluationAutomated Software Testing

  6. The Selection Rule Decides the Winner: A Pre-Registered Audit of Open-Set Graph Anomaly Detection

    Sep 27, 2026Farhan Shahriyar Hossain, Taufikur Rahman Fuad, Md Abrar Jahin +1Few-Shot Anomaly DetectionBenchmark Auditing

  7. Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark

    Sep 23, 2026Makar Ulesov, Vladislav Smirnov, Omar Ibrahim +1Benchmark AuditingLLM Auditing

  8. Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL

    Sep 22, 2026Jiamiao Liu, Dewen Qiao, Yu Zhang +1Probability CalibrationBenchmark Auditing

  9. ReSTI: A Source-Grounded Audit and Repair of STI-Bench

    Sep 21, 2026Pengzhan Sun, Ramanathan Rajaraman, Shiu-Hong Kao +2Spatial Reasoning BenchmarksSpatiotemporal Reasoning

  10. ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift

    Sep 20, 2026Kemal Kirtac, Carsten MapleDistribution Shift RobustnessBenchmark Auditing

  11. Auditing Bayesian Graph Alignment: Diagnostic Comparisons and Reference Failure

    Sep 19, 2026Melika Gorgi, Kourosh MirsohiMarkov Chain Monte CarloBenchmark Auditing

  12. DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?

    Sep 14, 2026Hongye Yang, Zhihao Xie, Shengjun Xiong +1Benchmark AuditingText-to-CAD Generation

  13. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

    Sep 10, 2026Koutian Wu, Junjie Zhou, Ergan Shang +6Benchmark AuditingBenchmark Design

  14. Stress-Testing Efficient Responsible-AI Evaluation: When Compute Savings Change Benchmark Conclusions

    Aug 31, 2026Ahmed El Kady, Aravind Narayanan, Rehana Noorani +2LLM EvaluationEnergy-Efficient ML

  15. Auditing MCQA Benchmarks through Probability Landscapes

    Aug 31, 2026Minsoo Song, Chanjun ParkBenchmark AuditingLLM Auditing

  16. CatchBench: When Can an Agent Failure Be Caught?

    Aug 24, 2026Yue Zhao, Mengyuan Li, Ruolin Li +5Agent Failure AnalysisBenchmark Auditing

  17. When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits

    Aug 8, 2026Ibne Farabi Shihab, Sanjeda Akter, Anuj SharmaStatistical Power AnalysisBenchmark Auditing

  18. On the missing benchmarks layer and a potential solution

    Aug 4, 2026Francis F Daniel, Mauro Ibañez, Francis Perelman +1LLM EvaluationBenchmark Auditing

  19. Coding Agents as Test-Suite Auditors: Finding What Official Suites Miss While Approaching What They Catch

    Aug 3, 2026Shuyang Xie, Shuxiao Xie, Feng Zhu +2Automated Software TestingBenchmark Auditing

  20. Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

    Jul 30, 2026Philipp D. Siedler, Jordan SassoonLLM EvaluationBenchmark Auditing