Benchmark Design

Latest papers 285

All topics
CardsList
  1. Benchmark Dataset for Catalysis on 2D MXenes

    May 30, 2026Pavlo Melnyk, Anmar Karmush, Mårten Wadenbäck +4Machine Learning Interatomic PotentialsBenchmark Design

  2. TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering

    May 29, 2026Jin Gao, Juntu Zhao, Zirui Zeng +5Protein Fitness PredictionBenchmark Design

  3. Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation

    May 29, 2026Andreas Haupt, Justin Hartenstein, Anka Reuel +2Algorithmic AuditingBenchmark Design

  4. CalArena: A Large-Scale Post-Hoc Calibration Benchmark

    May 28, 2026Eugène Berta, David Holzmüller, Francis Bach +1Post-Hoc CalibrationProbability Calibration

  5. CLUBench: A Clustering Benchmark

    May 28, 2026Feng Xiao, Dazhi Fu, Chris Ding +1Benchmark DesignClustering

  6. OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

    May 28, 2026Wanhao Liu, Jiaqing Xie, Qian Tan +10Benchmark DesignScientific QA

  7. Benchmarking AI for low-resource contexts: Thinking beyond leaderboards

    May 27, 2026Aakash Pant, Kavya Shah, Apoorv Agnihotri +3Benchmark DesignAutomated Evaluation

  8. BuddyBench: A Privacy-Constrained Multi-Task Benchmark for Pediatric Social-Communication Personalization

    May 27, 2026Jeyeon Eo, Joo Young Kim, Ran Ju +2Sequential RecommendationHealthcare

  9. Verifiable Benchmarking of Long-Horizon Spatial Biology

    May 27, 2026Ian Diks, Harihara Muralidharan, Tim Proctor +1Spatial TranscriptomicsBenchmark Design

  10. SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

    May 26, 2026Haosong Peng, Hao Li, Jiaqi Chen +10Representation LearningBenchmark Design

  11. GrowLoop: Self-Evolving Conversation Evaluation Seeded by Human

    May 26, 2026Yihang Lin, Yunze Gao, Zeyang Lin +3Benchmark DesignHuman-in-the-Loop Evaluation

  12. AssertLLM2: A Comprehensive LLM Benchmark for Assertion Generation from Design Specifications

    May 26, 2026Yuchao Wu, Wenji Fang, Jing Wang +3Formal Hardware VerificationLLM Evaluation

  13. Constraint acquisition needs better benchmarks

    May 25, 2026Rafał Stachowiak, Tomasz P. PawlakBenchmark ConstructionBenchmark Design

  14. Deployment-complete benchmarking

    May 25, 2026El Mustapha Mansouri, Keigo AraiBenchmark AuditingBenchmark Construction

  15. A Scalable Benchmark Test Suite for Dynamic Multi-Objective Optimization with a Changing Number of Objectives

    May 25, 2026Ke Shang, Zhiyun Xiao, Yuxuan Liu +3Benchmark DesignMulti-Objective Optimization

  16. Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

    May 25, 2026Zexin Zhuang, Yanhang Li, Zhichao FanLLM QuantizationStatistical Power Analysis

  17. SafetyRepro: Configuration-Conditional Rank Instability on Alignment Benchmarks

    May 25, 2026Yanhang Li, Zhichao Fan, Zexin ZhuangPairwise ComparisonPairwise Preference Evaluation

  18. ViroBench: Benchmarking Nucleotide Foundation Models on Viral Genomics Tasks

    May 25, 2026Dongxin Ye, Fang Hu, Han Hu +6Distribution ShiftBenchmark Design

  19. AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

    May 24, 2026Michael Hardy, Anka Reuel, Lijin Zhang +6LLM EvaluationBenchmark Design

  20. RealBench: Benchmarking Data-Driven Numerical Weather Forecasting Under Operational Conditions and Extreme Event Challenges

    May 24, 2026Ruize Li, Zhibin Wen, Tao Han +5Benchmark DesignForecasting Benchmarks

  21. MR-LiDAR: A Multi-Resolution Roadside LiDAR Benchmark for Perception Diagnostics and Deployment Guidance

    May 23, 2026Shunlai Cui, Peng Cao, Yuan Zhu +5Benchmark Design3D Object Detection

  22. How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness

    May 22, 2026Polina Gordienko, Georg Schollmeyer, Frauke Kreuter +1Benchmark Design

  23. Design and Report Benchmarks for Knowledge Work

    May 22, 2026Yining Hua, Hongbin Na, Cyrus Ayubcha +1Benchmark DesignAI Agent Evaluation