Benchmark Validity

Momentum

15 papers in the last four weeks, up 114% on the four weeks before. 0.1% of all new papers.

Jul 13Week of Sep 28

Latest papers 117

All topics
CardsList
  1. The Great Pretender: A Stochasticity Problem in LLM Jailbreak

    May 14, 2026Jean-Philippe Monteuuis, Cong Chen, Jonathan PetitLLM Safety BenchmarksLLM Evaluation

  2. Unsteady Metrics and Benchmarking Cultures of AI Model Builders

    May 13, 2026Stefan Baack, Christo Buschek, Maty BohacekLLM EvaluationBenchmark Design

  3. A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

    May 13, 2026Takehiro Ishikawa, Jon DukeDepression DetectionMental Health

  4. Log analysis is necessary for credible evaluation of AI agents

    May 8, 2026Peter Kirgis, Sayash Kapoor, Stephan Rabanser +8Agent Failure AnalysisAI Agent Evaluation

  5. Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking

    May 7, 2026Yang Xu, Jiefu Zhang, Haixiang Sun +3LLM EvaluationSelective Inference

  6. What Single-Prompt Accuracy Misses: A Multi-Variant Reliability Audit of Language Models

    May 3, 2026Ranit Karmakar, Jayita ChatterjeeLLM EvaluationLanguage Model Calibration

  7. BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

    Apr 27, 2026Xinming Tu, Tianze Wang, Yingzhou +4Benchmark AuditingAI Agent Benchmarks

  8. All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation

    Apr 27, 2026Leonardo Haw-Yang Foo, Chih-Kai Yang, Chen-An Li +2Audio-Language Model EvaluationAudio QA

  9. LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks

    Mar 20, 2026Xiang Long, Li Du, Yilong Xu +11Computer-Use Agent BenchmarksTool-Augmented Language Model Agents

  10. Unexplored flaws in multiple-choice VQA make benchmarking unreliable

    Nov 27, 2025Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf +3Visual Question AnsweringPrompt Sensitivity

  11. The Benchmarking Epistemology: Validity Theory for Evaluating Machine Learning Models

    Oct 27, 2025Timo Freiesleben, Sebastian ZezulkaBenchmark DesignBenchmark Validity

  12. fev-bench: A Realistic Benchmark for Time Series Forecasting

    Sep 30, 2025Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen +5Benchmark DesignTime Series Forecasting

  13. Who Benchmarks the Benchmarks? Towards Comprehensive Evaluation of Commonsense Reasoning Benchmarks

    Apr 10, 2025Pavel Chizhov, Anton Changalidis, Vishnu Prasad Vijaya Kumar +4Benchmark AuditingReasoning Evaluation

  14. Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models

    Mar 9, 2025Batu Guan, Xiao Wu, Yuanyuan Yuan +1Benchmark DesignData Contamination in Language Models

  15. GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods

    Feb 24, 2025Ruixuan Huang, Xunguang Wang, Zongjie Li +2Language Model Safety EvaluationLLM Jailbreak Attacks

  16. Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

    Date pendingWilliam CabanInter-Rater ReliabilityAI Agent Evaluation