Benchmark Design

Latest papers 282

All topics
CardsList
  1. TabJoinBench: A Benchmark for Joinable Table Discovery

    Sep 30, 2026Sandipan De, Jin Wang, Vivek GuptaBenchmark Design

  2. PhysicsMate: A Curriculum-Grounded Bengali Benchmark for Secondary Physics QA with Small-Model Adaptation

    Sep 30, 2026Rashid Azraf Jahin, Saadman Sajid, Khan Raiyan Ibne Reza +1Benchmark DesignScientific QA

  3. cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

    Sep 30, 2026Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck +3Computer-Use Agent BenchmarksBenchmark Design

  4. A 3GPP-Compliant Benchmark Dataset for RIS-Aided Beyond 5G Networks

    Sep 30, 2026Pujitha Mamillapalli, Pankaj Singh Rathour, Abhinav KumarReconfigurable Intelligent SurfacesBenchmark Design

  5. PDE-OBS: Controlled Evaluation Across Observation Patterns

    Sep 29, 2026Ruichen Xu, Siyao Wang, Fang Wan +7Benchmark DesignPDE Surrogate Modeling

  6. AraDynFact: Dynamic Evaluation of Factual Knowledge in Arabic

    Sep 28, 2026Ignacio Iacobacci, Faroq Altam, Zhaozhi Qian +1LLM EvaluationArabic NLP

  7. SymbolicArena: A Unified Infrastructure for Benchmark Distillation and Dynamic Evaluation in Symbolic Regression

    Sep 28, 2026Ziwen Zhang, Xiju Wu, Yuheng Jing +9Benchmark DesignAutomated Evaluation

  8. MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses

    Sep 27, 2026Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu +3LLM EvaluationBenchmark Design

  9. Artificial Societies Benchmark: A Validation Framework for Synthetic Research

    Sep 24, 2026Edoardo Chidichimo, Min Jun Jung, Felix P. S. Wallis +1Benchmark DesignComputational Social Science

  10. TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening

    Sep 24, 2026Surbhi Kumar, Yuhe Zhou, Varun Shiralkar +2Benchmark DesignVirtual Screening

  11. Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

    Sep 24, 2026Yezhou Cheng, Runjia Du, Zeming Liu +5Confidence Region EstimationUncertainty Quantification

  12. WPBench: A Comprehensive Benchmark for Wind Power Forecasting

    Sep 21, 2026Yuhan Zhu, Jilin Hu, Xinying Cai +7Multivariate Time Series ForecastingBenchmark Design

  13. VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks

    Sep 20, 2026Liyang Fan, Yingcheng Shi, Yongbin Li +7Benchmark DesignSoftware Engineering Benchmarks

  14. TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking

    Sep 16, 2026Giorgia Adorni, Michela Papandrea, Battista Rimoldi +1Text-to-Music GenerationMusic Generation Evaluation

  15. Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors

    Sep 16, 2026Peixuan Hou, Bin Chen, Li He +4Benchmark DesignUser Preference Modeling

  16. HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT Applications

    Sep 14, 2026Inesh Chakrabarti, Zejun Xiong, Pragya Sharma +1Cyber-Physical SystemsEfficient Neural Network Inference

  17. IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

    Sep 12, 2026Kainan Zhou, Zhaoyi Li, Janet Sung +2LLM EvaluationBenchmark Design

  18. Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens

    Sep 12, 2026Carl Edwards, Edward De Brouwer, Xiner Li +5Benchmark DesignAdaptive Sampling

  19. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

    Sep 10, 2026Koutian Wu, Junjie Zhou, Ergan Shang +6Benchmark AuditingBenchmark Design

  20. PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us

    Sep 5, 2026Gal Sapir, Alon Diament, Adva Wolf +7LLM EvaluationBenchmark Design

  21. RealCADBench: Benchmarking Parametric CAD Modeling from Industrial Design Intents

    Sep 3, 2026JoyIndustrial VisCAD Team, Linxin Cai, Qiuhe Hong +10Benchmark DesignParametric CAD Modeling

  22. StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

    Sep 1, 2026Yinghao Chen, Zixi Chen, Bingxiang He +7Benchmark DesignLanguage Model Self-Improvement

  23. Toward Workflow-Aware Benchmarking for Healthcare NLP Agents

    Aug 31, 2026Junyi Yao, Baichuan Li, Zihao Zheng +1Benchmark DesignLLM Agent Evaluation

  24. A Browser-Native Digital Test Range for Benchmarking 4D Ocean-Glider Planning Algorithms

    Aug 13, 2026Edward Holmberg, Elias Ioup, Mahdi AbdelguerfiBenchmark DesignUnderwater Robotics

  25. Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks

    Aug 13, 2026Wojciech Zarzecki, Jarosław ArabasEvolutionary OptimizationBenchmark Design

  26. LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

    Aug 13, 2026Chenrun Wang, Mingxuan Zhu, Tiancheng Huang +6LLM EvaluationBenchmark Design

  27. CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

    Aug 13, 2026Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto +3Benchmark DesignSynthetic Data Generation

  28. DiG-bench: Discovery in Games

    Aug 12, 2026Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13Benchmark DesignAI Agent Evaluation

  29. NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation

    Aug 12, 2026Jiarui Ma, Jianghan Wang, Yuheng Ma +2LLM EvaluationBenchmark Design

  30. How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models

    Aug 12, 2026Aleksandra Kalisz, Jack Simons, Krisztina Sinkovics +4Protein Structure PredictionInference-Time Optimization

  31. Large-scale AI-Ready Data for Anti-Cancer Drug Response Modeling

    Aug 11, 2026Vincent Lavelle, Yitan Zhu, Kaitlyn Marlor +2Benchmark DesignOOD Generalization

  32. IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

    Aug 5, 2026Shahd Gaben, Heba Sbahi, Samer Rashwani +5LLM EvaluationBenchmark Design

  33. Why Ranking Anomaly Detection Algorithms Isn't as Reliable as You May Think

    Aug 5, 2026Simon Klüttermann, Jérôme Rutinowski, Frederik Polachowski +1Benchmark DesignML Reproducibility

  34. OmniRouting: A Semantic-Coupled Multimodal Benchmark for Constraint-Aware Spatial Reasoning in PCB Routing

    Aug 5, 2026Taiting Lu, Kaiyuan Lin, Ziwei Dong +18Benchmark DesignElectronic Design Automation

  35. Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

    Aug 4, 2026William Bolton, Philip TorrAI Agents for Scientific DiscoveryBenchmark Design

  36. Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design

    Aug 4, 2026Zejun Liu, Jian Wu, Ru Peng +4AI for ScienceBenchmark Design

  37. FinVerse: Financial Time-Series Benchmark

    Aug 4, 2026Jaehoon Lee, Jun Seo, Seunghan Lee +9Benchmark DesignFinancial Forecasting

  38. On the missing benchmarks layer and a potential solution

    Aug 4, 2026Francis F Daniel, Mauro Ibañez, Francis Perelman +1LLM EvaluationBenchmark Auditing

  39. onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

    Aug 3, 2026Brandon Wang, Andrei S. Tyrin, Daniil A. BoikoLLM Safety BenchmarksLLM Evaluation

  40. From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution

    Aug 3, 2026Can Wang, Haoran Chen, Haowen Gao +3LLM EvaluationBenchmark Design

  41. How Benchmarks and Evaluation Protocols Shape Conclusions in Provenance-Based Intrusion Detection

    Aug 2, 2026Lorenzo Guerra, Thomas Chapuis, Guillaume Duc +2Data ProvenanceBenchmark Design