Benchmark Construction

Latest papers 51

All topics
CardsList
  1. WitnessGym: Benchmarking Coding Agents on the Construction of Bug Witnesses

    Sep 29, 2026Haomin Qi, Xiangzhe Xu, Yiming Huang +2Software Engineering AgentsBenchmark Construction

  2. TraceDance: An Automated System for Building Agent Behavior Benchmarks from Real-World Agent Deployment Traces

    Sep 27, 2026Dehai Min, Daoan Zhang, Yiming Zeng +13Benchmark ConstructionLLM Agent Evaluation

  3. A comparative assessment of global building and settlement datasets across geographic and settlement contexts

    Sep 23, 2026Rufai Omowunmi Balogun, Caroline Margaux Gevaert, Capucine Riom +5Benchmark ConstructionBuilding Footprint Extraction

  4. Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts

    Sep 16, 2026Guojun Zhu, Xunheng Huang, Peng Yin +3Counterfactual EvaluationBenchmark Construction

  5. Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice Evaluation

    Sep 14, 2026M. Ali BayramMultilingual Language Model EvaluationLLM Evaluation

  6. Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

    Sep 14, 2026Baoyang Jiang, Fengchun Zhang, Leyuan Wang +9Benchmark ConstructionEmbodied QA

  7. Field-level prediction of mid-plane stress tensor fields in concrete target penetration: a cross-velocity graph neural operator surrogate

    Sep 9, 2026Wenpu Du, Peng Zhou, Yunlong Xia +5Neural Surrogate ModelingBenchmark Construction

  8. A Large Open Multi-Energy Corpus of Soil Compaction Tests, with Machine-Learning Baselines

    Sep 3, 2026Sompote Youwai, Chana Phutthananon, Warat KongkitkulBenchmark ConstructionMaterials Property Prediction

  9. CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

    Aug 10, 2026Harmanjot Singh, Abhra Dubey, Jorge Alejandro Amador HerreraBenchmark ConstructionParametric CAD Modeling

  10. ArchEGraph: A Large-Scale Graph Dataset for Geometry-Topology-Physics Aligned Building Energy Modeling

    Aug 7, 2026Yihui Li, Yihui Chen, Kaidi Zha +7Benchmark ConstructionBuilding Energy Management

  11. Who Belongs in the Eval Set? A Capability-Taxonomy-Driven Pipeline for Curating Regression Eval Sets in Agent-Extensibility Platforms

    Aug 2, 2026Tezan Sahu, Aritra Das, Pankaj Mittal +1Benchmark ConstructionLLM Agent Evaluation

  12. Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation

    Jul 30, 2026Philipp D. Siedler, Jordan SassoonLLM EvaluationBenchmark Auditing

  13. SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining

    Jul 29, 2026Yi Cui, Zilin Wang, Yijie Xu +5Multimodal RobustnessBenchmark Construction

  14. HamQASBench: A Hamiltonian-Informed Diagnostic Benchmark for Evaluating Quantum Architecture Search

    Jul 6, 2026Jiayang Niu, Akib Karim, Yan Wang +5Benchmark ConstructionBenchmark Design

  15. MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

    Jul 2, 2026Yuanzhi Liu, Shousheng Zhao, Bo Zhou +2VLM EvaluationBenchmark Construction

  16. EEG Benchmarking Needs a Task Specification Layer: NeuroDoc for Rulebook-Guided, Executable Benchmark Construction

    Jun 22, 2026Chengxuan Qin, Zhige Chen, Shu Peng +9ElectroencephalographyBenchmark Construction

  17. CulMind: Benchmarking Multimodal Understanding and Reasoning in Chinese Cultural Heritage

    Jun 19, 2026Zhangwei Cao, Shuhan Fan, Yuting Wei +5Benchmark ConstructionReasoning Evaluation

  18. Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends

    Jun 10, 2026Jinshan Lai, Jianwei Hu, Baoyang Jiang +7Benchmark ConstructionBenchmark Design

  19. ClusBench: The Clustering Benchmark Data Resource You've All Been Waiting For (?)

    Jun 9, 2026David P. HofmeyrBenchmark ConstructionClustering

  20. Towards Event-Robust Acoustic Scene Classification

    Jun 5, 2026Yiqiang Cai, Bohan Hu, Yu Yang +3Benchmark ConstructionAudio Classification

  21. Brick-Composer: Using MLLMs for Assembly with Diverse Bricks

    Jun 3, 2026Jiateng Liu, Bingxuan Li, Zhenhailong Wang +8Visual Spatial ReasoningBenchmark Construction

  22. SC3: The Multi-Solvent Solubility Challenge and Benchmark

    Jun 3, 2026Vansh Ramani, Har Ashish Arora, Dhairya Kuchhal +4Benchmark ConstructionBenchmark Validity

  23. Position: Every Ground Truth is a Human Construction, not an Objective Truth

    May 28, 2026Charlotte Högberg, Ericka Johnson, Kiri L. WagstaffBenchmark Construction

  24. A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

    May 27, 2026Tomer Keren, Nitay Calderon, Asaf Yehudai +3Computer-Use Agent BenchmarksTool-Use Evaluation

  25. ATLAS: All-round Testing of Long-context Abilities across Scales

    May 27, 2026Deli Huang, Cunguang Wang, Hongyin Tang +15LLM EvaluationBenchmark Construction

  26. Constraint acquisition needs better benchmarks

    May 25, 2026Rafał Stachowiak, Tomasz P. PawlakBenchmark ConstructionBenchmark Design