Artificial Intelligence Benchmarks

Latest papers 41

All topics
CardsList
  1. The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

    Sep 10, 2026Makoto Fukushima, Hua-Dong Xiong, Ehsan Moradi PariAgentic EvaluationsArtificial Intelligence Benchmarks

  2. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

    Sep 10, 2026Koutian Wu, Junjie Zhou, Ergan Shang +6Artificial Intelligence BenchmarksLarge Language Model Benchmarks

  3. DiG-bench: Discovery in Games

    Aug 12, 2026Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13Agentic BenchmarksArtificial Intelligence Benchmarks

  4. LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

    Aug 10, 2026Chenyang Li, Zejia Feng, Yuqin Huang +2Legal Reasoning TasksLegal Domain

  5. Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

    Aug 4, 2026William Bolton, Philip TorrScientific IdeationArtificial Intelligence Benchmarks

  6. On the missing benchmarks layer and a potential solution

    Aug 4, 2026Francis F Daniel, Mauro Ibañez, Francis Perelman +1Artificial Intelligence BenchmarksArtificial Intelligence Systems

  7. On the missing data layer and a potential solution

    Aug 3, 2026Francis F Daniel, Mauro Ibañez, Francis Perelman +1Artificial Intelligence InfrastructureLarge Databases

  8. Evaluating RAG for French immigration law: a benchmark and baseline study

    Jul 27, 2026Annia Abtout, Julien Delaunay, Monika Ewa RakoczyLuxembourgishArtificial Intelligence Benchmarks

  9. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

    Jul 17, 2026Ajay Patel, Kartik Hosanagar, Ramayya Krishnan +2Artificial Intelligence BenchmarksHuman-Annotated Benchmark

  10. Can We Trust Item Response Theory for AI Evaluation?

    Jul 16, 2026Han Jiang, Sunbeom Kwon, Jinwen Luo +2Artificial Intelligence BenchmarksItem Response Theory

  11. MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection

    Jul 16, 2026Goktug OzkanArtificial Intelligence SafetyClinician Trust

  12. Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models

    Jul 9, 2026Matteo Santelmo, Xiuying Wei, Israa Fakih +5Artificial Intelligence BenchmarksArtificial Intelligence Models

  13. Two AI Metrics Diverged: Will it Make All the Difference?

    Jul 1, 2026Alex Fogelson, Zachary A. Brown, Hans Gundlach +2Artificial Intelligence BenchmarksPerformance Metric

  14. BenchX: Benchmarking AI Models for Cancer Detection and Localization with Demographic and Protocol Biases

    Jun 23, 2026Qi Chen, Wenxuan Li, Pedro R. A. S. Bassi +14Cancer DetectionArtificial Intelligence Benchmarks

  15. ForecastBench-Sim: A Simulated-World Forecasting Benchmark

    Jun 17, 2026Jaeho Lee, Nick Merrill, Ezra KargerArtificial Intelligence Benchmarks

  16. Bounded Difference Concentration for Infinitely Exchangeable Sequences with Applications to AI Benchmark Uncertainty

    Jun 16, 2026Fangyuan Lin, Spencer Frei, Victor H. de la PenaConcentration BoundsArtificial Intelligence Benchmarks

  17. Automated Benchmark Auditing for AI Agents and Large Language Models

    May 25, 2026Junlin Wang, Federico Bianchi, Shang Zhu +4Artificial Intelligence BenchmarksAgentic Benchmarks

  18. AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

    May 24, 2026Michael Hardy, Anka Reuel, Lijin Zhang +6Artificial Intelligence BenchmarksLeaderboard

  19. RealBench: Benchmarking Data-Driven Numerical Weather Forecasting Under Operational Conditions and Extreme Event Challenges

    May 24, 2026Ruize Li, Zhibin Wen, Tao Han +5Artificial Intelligence Weather ModelsNumerical Weather Prediction

  20. Position: State-of-the-Art Claims Require State-of-the-Art Evidence

    May 17, 2026YongKyung OhArtificial Intelligence BenchmarksEvidence Conflict

  21. Unsteady Metrics and Benchmarking Cultures of AI Model Builders

    May 13, 2026Stefan Baack, Christo Buschek, Maty BohacekArtificial Intelligence BenchmarksMulti-Dimensional Evaluation

  22. NeuralBench: A Unifying Framework to Benchmark NeuroAI Models

    May 8, 2026Hubert Banville, Stéphane d'Ascoli, Simon Dahan +12Small-Sample Electroencephalogram DatasetsArtificial Intelligence Benchmarks

  23. Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference

    May 1, 2026Yuxuan Gao, Megan Wang, Yi Ling YuArtificial Intelligence BenchmarksArtificial Intelligence Systems

  24. Railway Artificial Intelligence Learning Benchmark (RAIL-BENCH): A Benchmark Suite for Perception in the Railway Domain

    Apr 24, 2026Annika Bätz, Pavel Klasek, Seo-Young Ham +3RelightingAutonomous Driving Perception

  25. E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes

    Apr 20, 2026Koya Sakamoto, Taiki Miyanishi, Daichi Azuma +6Viewpoint-Dependent Active PerceptionMultiple Vision Tasks

  26. MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition

    Apr 17, 2026Farhad Abtahi, Abdolamir Karbalaie, Eduardo Illueca-Fernandez +1Artificial Intelligence BenchmarksMetacognition

  27. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

    Feb 18, 2026Mubashara Akhtar, Anka Reuel, Prajna Soni +34Artificial Intelligence BenchmarksLarge Language Model Benchmarks

  28. BRIDGE: Predicting Human Task Completion Time From Model Performance

    Feb 6, 2026Fengyuan Liu, Jay Gala, Nilaksh +3Artificial Intelligence BenchmarksPsychometric Properties