Benchmark Validity

Momentum

15 papers in the last four weeks, up 114% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 117

All topics
CardsList
  1. The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

    Jun 2, 2026Wojciech Zarzecki, Jan Dubiński, Sebastian CygertBenchmark AuditingLLM Auditing

  2. Consistency evaluation of benchmarks used for causal discovery

    Jun 1, 2026Yuzhe Zhang, Chihui Chen, Lina Yao +1Causal DiscoveryBenchmark Validity

  3. Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation

    May 29, 2026Andreas Haupt, Justin Hartenstein, Anka Reuel +2Algorithmic AuditingBenchmark Design

  4. Auditing LLM Benchmarks with Item Response Theory

    May 28, 2026Sander Land, Daniel M. BikelBenchmark AuditingLLM Auditing

  5. NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation Models

    May 28, 2026Anany KotawalaData LeakageLLM Evaluation

  6. FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

    May 27, 2026Nishal Thomas, Noel ThomasMathematical Reasoning BenchmarksBenchmark Auditing

  7. The Trust Paradox: How CS Researchers Engage LLM Leaderboards

    May 27, 2026Pouya Sadeghi, Anamaria Crisan, Jimmy LinModel SelectionLLM Evaluation

  8. Code as a Weapon: A Consensus-Labeled Prompt Bank for Measuring Coding-Model Compliance with Malicious-Code Requests

    May 27, 2026Richard J. Young, Gregory D. MoodyClassificationLLM Safety Evaluation

  9. A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test

    May 27, 2026Camilo Chacón Sartori, José H. GarcíaRetrieval-Augmented GenerationLLM-as-a-Judge

  10. Automated Benchmark Auditing for AI Agents and Large Language Models

    May 25, 2026Junlin Wang, Federico Bianchi, Shang Zhu +4LLM EvaluationBenchmark Auditing

  11. Deployment-complete benchmarking

    May 25, 2026El Mustapha Mansouri, Keigo AraiBenchmark AuditingBenchmark Construction

  12. Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

    May 25, 2026Zexin Zhuang, Yanhang Li, Zhichao FanLLM QuantizationStatistical Power Analysis

  13. SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks

    May 25, 2026Yanhang Li, Zhichao Fan, Zexin ZhuangPairwise ComparisonPairwise Preference Evaluation

  14. AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

    May 24, 2026Michael Hardy, Anka Reuel, Lijin Zhang +6LLM EvaluationBenchmark Design

  15. Design and Report Benchmarks for Knowledge Work

    May 22, 2026Yining Hua, Hongbin Na, Cyrus Ayubcha +1Benchmark DesignAI Agent Evaluation

  16. Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

    May 22, 2026Chuyifei Zhang, Hongyu Cui, Xiaowen Huang +1LLM EvaluationLong-Context Language Model Evaluation

  17. Decomposing and Measuring Evaluation Awareness

    May 21, 2026Changling Li, Terry Jingchen Zhang, Jie Zhang +4LLM EvaluationLanguage Model Safety Evaluation

  18. Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

    May 21, 2026Naveen Raman, Santiago Cortes-Gomez, Mateo Dulce Rubio +2HealthcareBenchmark Design

  19. Measuring Security Without Fooling Ourselves: Why Benchmarking Agents Is Hard

    May 21, 2026Sahar Abdelnabi, Chris Hicks, Konrad Rieck +1AI Agent EvaluationAI Agent Security

  20. Provable Joint Decontamination for Benchmarking Multiple Large Language Models

    May 20, 2026Zhenlong Liu, Hao Zeng, Hongxin WeiLLM EvaluationBenchmark Contamination

  21. AgentAtlas: Beyond Outcome Leaderboards for LLM Agents

    May 19, 2026Parsa Mazaheri, Kasra MazaheriAgent Failure AnalysisComputer-Use Agent Benchmarks

  22. LLM Benchmark Datasets Should Be Contamination-Resistant

    May 19, 2026Ali Al-Lawati, Jason Lucas, Dongwon Lee +1LLM EvaluationBenchmark Contamination

  23. Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks

    May 18, 2026Yubin Qu, Ying Zhang, Yanjun Zhang +4Computer-Use Agent BenchmarksCoding Agents

  24. Position: State-of-the-Art Claims Require State-of-the-Art Evidence

    May 17, 2026YongKyung OhBenchmark Validity

  25. The Great Pretender: A Stochasticity Problem in LLM Jailbreak

    May 14, 2026Jean-Philippe Monteuuis, Cong Chen, Jonathan PetitLLM Safety BenchmarksLLM Evaluation

  26. Unsteady Metrics and Benchmarking Cultures of AI Model Builders

    May 13, 2026Stefan Baack, Christo Buschek, Maty BohacekLLM EvaluationBenchmark Design

  27. A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

    May 13, 2026Takehiro Ishikawa, Jon DukeDepression DetectionMental Health

  28. Log analysis is necessary for credible evaluation of AI agents

    May 8, 2026Peter Kirgis, Sayash Kapoor, Stephan Rabanser +8Agent Failure AnalysisAI Agent Evaluation

  29. Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking

    May 7, 2026Yang Xu, Jiefu Zhang, Haixiang Sun +3LLM EvaluationSelective Inference

  30. What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models

    May 3, 2026Ranit Karmakar, Jayita ChatterjeeLLM EvaluationLanguage Model Calibration

  31. BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

    Apr 27, 2026Xinming Tu, Tianze Wang, Yingzhou +4Benchmark AuditingAI Agent Benchmarks

  32. All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation

    Apr 27, 2026Leonardo Haw-Yang Foo, Chih-Kai Yang, Chen-An Li +2Audio-Language Model EvaluationAudio QA

  33. LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

    Mar 20, 2026Xiang Long, Li Du, Yilong Xu +11Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  34. Unexplored flaws in multiple-choice VQA make benchmarking unreliable

    Nov 27, 2025Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf +3Visual Question AnsweringPrompt Sensitivity

  35. The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models

    Oct 27, 2025Timo Freiesleben, Sebastian ZezulkaBenchmark DesignBenchmark Validity

  36. fev-bench: A Realistic Benchmark for Time Series Forecasting

    Sep 30, 2025Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen +5Benchmark DesignTime Series Forecasting

  37. Who Benchmarks the Benchmarks? Towards Comprehensive Evaluation of Commonsense Reasoning Benchmarks

    Apr 10, 2025Pavel Chizhov, Anton Changalidis, Vishnu Prasad Vijaya Kumar +4Benchmark AuditingReasoning Evaluation

  38. Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models

    Mar 9, 2025Batu Guan, Xiao Wu, Yuanyuan Yuan +1Benchmark DesignData Contamination in Language Models

  39. GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods

    Feb 24, 2025Ruixuan Huang, Xunguang Wang, Zongjie Li +2Language Model Safety EvaluationLLM Jailbreak Attacks

  40. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

    Date pendingWilliam CabanInter-Rater ReliabilityAI Agent Evaluation