Benchmark Design

Latest papers 285

All topics
CardsList
  1. CANDO: Cooperative Agentic Network for Layout Design Optimization

    Oct 7, 2026Athanasios Masouris, Zheng Jing, Benjamin Sam Chandler +1Benchmark DesignCombinatorial Optimization

  2. The AI Evaluation Ecosystem

    Oct 7, 2026Yash Dave, Sang T. Truong, Serena Wang +1AI Agent EvaluationBenchmark Design

  3. TabJoinBench: A Benchmark for Joinable Table Discovery

    Sep 30, 2026Sandipan De, Jin Wang, Vivek GuptaBenchmark Design

  4. PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation

    Sep 30, 2026Rashid Azraf Jahin, Saadman Sajid, Khan Raiyan Ibne Reza +1Benchmark DesignScientific QA

  5. cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

    Sep 30, 2026Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck +3Computer-Use Agent BenchmarksBenchmark Design

  6. A 3GPP-Compliant Benchmark Dataset for RIS-Aided Beyond 5G Networks

    Sep 30, 2026Pujitha Mamillapalli, Pankaj Singh Rathour, Abhinav KumarReconfigurable Intelligent SurfacesBenchmark Design

  7. PDE-OBS: Controlled Evaluation Across Observation Patterns

    Sep 29, 2026Ruichen Xu, Siyao Wang, Fang Wan +7Benchmark DesignPDE Surrogate Modeling

  8. AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic

    Sep 28, 2026Ignacio Iacobacci, Faroq Altam, Zhaozhi Qian +1LLM EvaluationArabic NLP

  9. SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression

    Sep 28, 2026Ziwen Zhang, Xiju Wu, Yuheng Jing +9Benchmark DesignAutomated Evaluation

  10. MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses

    Sep 27, 2026Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu +3LLM EvaluationBenchmark Design

  11. Artificial Societies Benchmark: A Validation Framework for Synthetic Research

    Sep 24, 2026Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis +1Benchmark DesignComputational Social Science

  12. TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

    Sep 24, 2026Surbhi Kumar, Yuhe Zhou, Varun Shiralkar +2Benchmark DesignVirtual Screening

  13. Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

    Sep 24, 2026Yezhou Cheng, Runjia Du, Zeming Liu +5Confidence Region EstimationUncertainty Quantification

  14. WPBench: A Comprehensive Benchmark for Wind Power Forecasting

    Sep 21, 2026Yuhan Zhu, Jilin Hu, Xinying Cai +7Multivariate Time Series ForecastingBenchmark Design

  15. VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

    Sep 20, 2026Liyang Fan, Yingcheng Shi, Yongbin Li +7Benchmark DesignSoftware Engineering Benchmarks

  16. TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking

    Sep 16, 2026Giorgia Adorni, Michela Papandrea, Battista Rimoldi +1Text-to-Music GenerationMusic Generation Evaluation

  17. Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors

    Sep 16, 2026Peixuan Hou, Bin Chen, Li He +4Benchmark DesignUser Preference Modeling

  18. HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT Applications

    Sep 14, 2026Inesh Chakrabarti, Zejun Xiong, Pragya Sharma +1Cyber-Physical SystemsEfficient Neural Network Inference

  19. IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

    Sep 12, 2026Kainan Zhou, Zhaoyi Li, Janet Sung +2LLM EvaluationBenchmark Design

  20. Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens

    Sep 12, 2026Carl Edwards, Edward De Brouwer, Xiner Li +5Benchmark DesignAdaptive Sampling

  21. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

    Sep 10, 2026Koutian Wu, Junjie Zhou, Ergan Shang +6Benchmark AuditingBenchmark Design

  22. PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us

    Sep 5, 2026Gal Sapir, Alon Diament, Adva Wolf +7LLM EvaluationBenchmark Design

  23. RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

    Sep 3, 2026JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong +10Benchmark DesignParametric CAD Modeling

  24. StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

    Sep 1, 2026Yinghao Chen, Zixi Chen, Bingxiang He +7Benchmark DesignLanguage Model Self-Improvement

  25. Toward Workflow-Aware Benchmarking for Healthcare NLP Agents

    Aug 31, 2026Junyi Yao, Baichuan Li, Zihao Zheng +1Benchmark DesignLLM Agent Evaluation

  26. A Browser-Native Digital Test Range for Benchmarking 4D Ocean-Glider Planning Algorithms

    Aug 13, 2026Edward Holmberg, Elias Ioup, Mahdi AbdelguerfiBenchmark DesignUnderwater Robotics

  27. Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks

    Aug 13, 2026Wojciech Zarzecki, Jarosław ArabasEvolutionary OptimizationBenchmark Design

  28. LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

    Aug 13, 2026Chenrun Wang, Mingxuan Zhu, Tiancheng Huang +6LLM EvaluationBenchmark Design

  29. CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

    Aug 13, 2026Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto +3Benchmark DesignSynthetic Data Generation

  30. DiG-bench: Discovery in Games

    Aug 12, 2026Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13Benchmark DesignAI Agent Evaluation

  31. NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation

    Aug 12, 2026Jiarui Ma, Jianghan Wang, Yuheng Ma +2LLM EvaluationBenchmark Design

  32. How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models

    Aug 12, 2026Aleksandra Kalisz, Jack Simons, Krisztina Sinkovics +4Protein Structure PredictionInference-Time Optimization

  33. Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

    Aug 11, 2026Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor +2Benchmark DesignOOD Generalization

  34. IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

    Aug 5, 2026Shahd Gaben, Heba Sbahi, Samer Rashwani +5LLM EvaluationBenchmark Design

  35. Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think

    Aug 5, 2026Simon Klüttermann, Jérôme Rutinowski, Frederik Polachowski +1Benchmark DesignML Reproducibility

  36. OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

    Aug 5, 2026Taiting Lu, Kaiyuan Lin, Ziwei Dong +18Benchmark DesignElectronic Design Automation

  37. Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

    Aug 4, 2026William Bolton, Philip TorrAI Agents for Scientific DiscoveryBenchmark Design

  38. Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

    Aug 4, 2026Zejun Liu, Jian Wu, Ru Peng +4AI for ScienceBenchmark Design

  39. FinVerse: Financial Time-Series Benchmark

    Aug 4, 2026Jaehoon Lee, Jun Seo, Seunghan Lee +9Benchmark DesignFinancial Forecasting

  40. On the missing benchmarks layer and a potential solution

    Aug 4, 2026Francis F Daniel, Mauro Ibañez, Francis Perelman +1LLM EvaluationBenchmark Auditing

  41. onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

    Aug 3, 2026Brandon Wang, Andrei S. Tyrin, Daniil A. BoikoLLM Safety BenchmarksLLM Evaluation

  42. From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

    Aug 3, 2026Can Wang, Haoran Chen, Haowen Gao +3LLM EvaluationBenchmark Design

  43. How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection

    Aug 2, 2026Lorenzo Guerra, Thomas Chapuis, Guillaume Duc +2Data ProvenanceBenchmark Design