Benchmark Design

Latest papers 285

All topics
CardsList
  1. PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us

    Sep 5, 2026Gal Sapir, Alon Diament, Adva Wolf +7LLM EvaluationBenchmark Design

  2. RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

    Sep 3, 2026JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong +10Benchmark DesignParametric CAD Modeling

  3. StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

    Sep 1, 2026Yinghao Chen, Zixi Chen, Bingxiang He +7Benchmark DesignLanguage Model Self-Improvement

  4. Toward Workflow-Aware Benchmarking for Healthcare NLP Agents

    Aug 31, 2026Junyi Yao, Baichuan Li, Zihao Zheng +1Benchmark DesignLLM Agent Evaluation

  5. A Browser-Native Digital Test Range for Benchmarking 4D Ocean-Glider Planning Algorithms

    Aug 13, 2026Edward Holmberg, Elias Ioup, Mahdi AbdelguerfiBenchmark DesignUnderwater Robotics

  6. Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks

    Aug 13, 2026Wojciech Zarzecki, Jarosław ArabasEvolutionary OptimizationBenchmark Design

  7. LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

    Aug 13, 2026Chenrun Wang, Mingxuan Zhu, Tiancheng Huang +6LLM EvaluationBenchmark Design

  8. CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

    Aug 13, 2026Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto +3Benchmark DesignSynthetic Data Generation

  9. DiG-bench: Discovery in Games

    Aug 12, 2026Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13Benchmark DesignAI Agent Evaluation

  10. NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation

    Aug 12, 2026Jiarui Ma, Jianghan Wang, Yuheng Ma +2LLM EvaluationBenchmark Design

  11. How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models

    Aug 12, 2026Aleksandra Kalisz, Jack Simons, Krisztina Sinkovics +4Protein Structure PredictionInference-Time Optimization

  12. Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

    Aug 11, 2026Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor +2Benchmark DesignOOD Generalization

  13. IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

    Aug 5, 2026Shahd Gaben, Heba Sbahi, Samer Rashwani +5LLM EvaluationBenchmark Design

  14. Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think

    Aug 5, 2026Simon Klüttermann, Jérôme Rutinowski, Frederik Polachowski +1Benchmark DesignML Reproducibility

  15. OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

    Aug 5, 2026Taiting Lu, Kaiyuan Lin, Ziwei Dong +18Benchmark DesignElectronic Design Automation

  16. Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

    Aug 4, 2026William Bolton, Philip TorrAI Agents for Scientific DiscoveryBenchmark Design

  17. Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

    Aug 4, 2026Zejun Liu, Jian Wu, Ru Peng +4AI for ScienceBenchmark Design

  18. FinVerse: Financial Time-Series Benchmark

    Aug 4, 2026Jaehoon Lee, Jun Seo, Seunghan Lee +9Benchmark DesignFinancial Forecasting

  19. On the missing benchmarks layer and a potential solution

    Aug 4, 2026Francis F Daniel, Mauro Ibañez, Francis Perelman +1LLM EvaluationBenchmark Auditing

  20. onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

    Aug 3, 2026Brandon Wang, Andrei S. Tyrin, Daniil A. BoikoLLM Safety BenchmarksLLM Evaluation

  21. From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

    Aug 3, 2026Can Wang, Haoran Chen, Haowen Gao +3LLM EvaluationBenchmark Design

  22. How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection

    Aug 2, 2026Lorenzo Guerra, Thomas Chapuis, Guillaume Duc +2Data ProvenanceBenchmark Design