Benchmark Validity

Momentum

15 papers in the last four weeks, up 114% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 117

All topics
CardsList
  1. The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

    Jun 2, 2026Wojciech Zarzecki, Jan Dubiński, Sebastian CygertBenchmark AuditingLLM Auditing

  2. Consistency evaluation of benchmarks used for causal discovery

    Jun 1, 2026Yuzhe Zhang, Chihui Chen, Lina Yao +1Causal DiscoveryBenchmark Validity

  3. Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation

    May 29, 2026Andreas Haupt, Justin Hartenstein, Anka Reuel +2Algorithmic AuditingBenchmark Design

  4. Auditing LLM Benchmarks with Item Response Theory

    May 28, 2026Sander Land, Daniel M. BikelBenchmark AuditingLLM Auditing

  5. NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation Models

    May 28, 2026Anany KotawalaData LeakageLLM Evaluation

  6. FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

    May 27, 2026Nishal Thomas, Noel ThomasMathematical Reasoning BenchmarksBenchmark Auditing

  7. The Trust Paradox: How CS Researchers Engage LLM Leaderboards

    May 27, 2026Pouya Sadeghi, Anamaria Crisan, Jimmy LinModel SelectionLLM Evaluation

  8. Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

    May 27, 2026Richard J. Young, Gregory D. MoodyClassificationLLM Safety Evaluation

  9. A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test

    May 27, 2026Camilo Chacón Sartori, José H. GarcíaRetrieval-Augmented GenerationLLM-as-a-Judge

  10. Automated Benchmark Auditing for AI Agents and Large Language Models

    May 25, 2026Junlin Wang, Federico Bianchi, Shang Zhu +4LLM EvaluationBenchmark Auditing

  11. Deployment-complete benchmarking

    May 25, 2026El Mustapha Mansouri, Keigo AraiBenchmark AuditingBenchmark Construction

  12. Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

    May 25, 2026Zexin Zhuang, Yanhang Li, Zhichao FanLLM QuantizationStatistical Power Analysis

  13. SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks

    May 25, 2026Yanhang Li, Zhichao Fan, Zexin ZhuangPairwise ComparisonPairwise Preference Evaluation

  14. AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

    May 24, 2026Michael Hardy, Anka Reuel, Lijin Zhang +6LLM EvaluationBenchmark Design

  15. Design and Report Benchmarks for Knowledge Work

    May 22, 2026Yining Hua, Hongbin Na, Cyrus Ayubcha +1Benchmark DesignAI Agent Evaluation

  16. Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

    May 22, 2026Chuyifei Zhang, Hongyu Cui, Xiaowen Huang +1LLM EvaluationLong-Context Language Model Evaluation

  17. Decomposing and Measuring Evaluation Awareness

    May 21, 2026Changling Li, Terry Jingchen Zhang, Jie Zhang +4LLM EvaluationLanguage Model Safety Evaluation

  18. Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

    May 21, 2026Naveen Raman, Santiago Cortes-Gomez, Mateo Dulce Rubio +2HealthcareBenchmark Design

  19. Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard

    May 21, 2026Sahar Abdelnabi, Chris Hicks, Konrad Rieck +1AI Agent EvaluationAI Agent Security

  20. Provable Joint Decontamination for Benchmarking Multiple Large Language Models

    May 20, 2026Zhenlong Liu, Hao Zeng, Hongxin WeiLLM EvaluationBenchmark Contamination

  21. AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

    May 19, 2026Parsa Mazaheri, Kasra MazaheriAgent Failure AnalysisComputer-Use Agent Benchmarks

  22. LLM Benchmark Datasets Should Be Contamination-Resistant

    May 19, 2026Ali Al-Lawati, Jason Lucas, Dongwon Lee +1LLM EvaluationBenchmark Contamination

  23. Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks

    May 18, 2026Yubin Qu, Ying Zhang, Yanjun Zhang +4Computer-Use Agent BenchmarksCoding Agents

  24. Position: State-of-the-Art Claims Require State-of-the-Art Evidence

    May 17, 2026YongKyung OhBenchmark Validity