Benchmark Design

Latest papers 285

All topics
CardsList
  1. PSI-Bench: Interpretable and Clinically Meaningful Evaluation of Depression Patient Simulators

    Apr 28, 2026Nguyen Khoi Hoang, Shuhaib Mehri, Tse-An Hsu +4Benchmark DesignClinical Language Model Evaluation

  2. TrialCalibre: A Fully Automated Causal Engine for RCT Benchmarking and Observational Trial Calibration

    Apr 28, 2026Amir Habibdoust, Xing SongCausal Effect EstimationBenchmark Design

  3. EOS-Bench: A Comprehensive Benchmark for Earth Observation Satellite Scheduling

    Apr 28, 2026Qian Yin, Jiaxing Li, Jiaqi Cheng +23Benchmark ConstructionBenchmark Design

  4. Benchmarking bandgap prediction in semiconductors under experimental and realistic evaluation settings

    Apr 28, 2026Haolin Wang, Xianyuan Liu, Anna Jungbluth +3Benchmark DesignInterpretable ML

  5. Benchmarking Stopping Criteria for Evolutionary Multi-objective Optimization

    Apr 28, 2026Kenji Kitamura, Ryoji TanabeBenchmark DesignMulti-Objective Evolutionary Optimization

  6. Energy-Arena: A Dynamic Benchmark for Operational Energy Forecasting

    Apr 27, 2026Max Kleinebrahm, Jonathan Berrisch, Philipp Eiser +11Benchmark DesignTime Series Forecasting

  7. On Benchmark Hacking in ML Contests: Modeling, Insights and Design

    Apr 24, 2026Xiaoyun Qiu, Yang Yu, Haifeng XuBenchmark Design

  8. OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving

    Apr 23, 2026Xinyu Zhang, Boxuan Zhang, Yuchen Wan +5LLM EvaluationBenchmark Design

  9. Performance Anomaly Detection in Athletics: A Benchmarking System with Visual Analytics

    Apr 23, 2026Blessed Madukoma, Prasenjit MitraBenchmark DesignTime-Series Anomaly Detection

  10. QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks

    Apr 20, 2026Taylor Lundy, Narun K. Raman, Kevin Leyton-BrownLLM EvaluationBayesian Optimization

  11. In Search of Lost DNA Sequence Pretraining

    Apr 17, 2026Zhijiang Tang, Jiaxin Qi, Yan Cui +3Language Model PretrainingMasked Language Modeling

  12. MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events

    Apr 16, 2026Raunak Agarwal, Markus Wenzel, Simon Baur +3Multi-Label ClassificationHealthcare

  13. An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics

    Apr 16, 2026Miri Liu, ChengXiang ZhaiAI for ScienceBenchmark Design

  14. FedGUI: Benchmarking Federated GUI Agents across Heterogeneous Platforms, Devices, and Operating Systems

    Apr 16, 2026Wenhao Wang, Haoting Shi, Mengying Yuan +7Benchmark DesignGUI Agents

  15. NanoBench: A Multi-Task Benchmark Dataset for Nano-Quadrotor System Identification, Control, and State Estimation

    Mar 10, 2026Syed Izzat Ullah, Jose BacaBenchmark DesignNonlinear System Identification

  16. PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment

    Mar 9, 2026Bence Mark Halpern, Thomas Tienkamp, Defne Abur +1Benchmark DesignSpeech Quality Assessment

  17. FlexMS: A Unified Public Benchmark for Molecule Tandem Mass Spectrum Prediction

    Feb 26, 2026Yunhua Zhong, Yixuan Tang, Yifan Li +5Benchmark DesignMass Spectrometry

  18. Two-Bridge: Exclusive Objectives and Extended Horizon StarCraft II Benchmark

    Feb 19, 2026Sourav Panda, Tanmay Ambadkar, Shreyash Kale +2RL BenchmarksGame-Playing Agents

  19. When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

    Feb 18, 2026Mubashara Akhtar, Anka Reuel, Prajna Soni +34LLM EvaluationBenchmark Design

  20. An Evolutionary Framework for Automatic Optimization Benchmark Generation via Large Language Models

    Jan 19, 2026Yuhiro Ono, Tomohiro Harada, Yukiya MiuraEvolutionary OptimizationBenchmark Design

  21. DR-Arena: an Automated Evaluation Framework for Deep Research Agents

    Jan 15, 2026Yiwen Gao, Ruochen Zhao, Yang Deng +1Deep Research AgentsBenchmark Design

  22. Large-scale benchmarking of multi-objective soft-computing metaheuristics for redundancy allocation in repairable k-out-of-n systems

    Dec 20, 2025Mateusz Oszczypała, David Ibehej, Jakub KudelaBenchmark DesignMetaheuristic Optimization

  23. SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?

    Nov 8, 2025Jeffrey Jian Ma, Milad Hashemi, Amir Yazdanbakhsh +5Software Engineering AgentsBenchmark Design

  24. MLPerf Automotive

    Oct 31, 2025Radoyeh Shojaei, Predrag Djurdjevic, Mostafa El-Khamy +7End-to-End Autonomous DrivingBenchmark Design