cs.AISep 27, 2026

MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses

Authors: Xuanjun Chen, Hua-Hsuan Chen, Wei-Chung Lu, Yinghao Ma, Jyh-Shing Roger Jang, Hung-yi Lee

Organizations: National Taiwan University · NTU AI-Core · Queen Mary University of London

Abstract

Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving beyond the evolution of isolated tasks, we propose to optimize the benchmark generation workflow itself end to end with MetaBench-Harness, a dual-loop search framework. Specifically, the inner loop utilizes a benchmark harness to generate a new benchmark in each round, while the outer meta-harness orchestration layer iteratively refines and searches over harness implementations based on historical evolution trajectories. By applying MetaBench-Harness to the competitive programming CodeContests and Olympiad mathematics AIME-2024 datasets, we demonstrate that the evolved benchmarks are challenging and discriminative for frontier models. Trajectory and quality analyses verify that MetaBench-Harness enables multi-dimensional evolution, steadily improving evolution reasonableness, benchmark competency, and evaluator robustness across successive rounds. Furthermore, case studies reveal its effective utilization of diverse difficulty levers to reframe problems and elevate required capabilities. Ultimately, this work provides a solution to the pressing challenge of benchmark saturation.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

    May 31, 2026Yangzhen Wu, Aaron J. Li, Wenjie Ma +10Large Language Model BenchmarksSynthetic Task

  2. An Evolutionary Framework for Automatic Optimization Benchmark Generation via Large Language Models

    Jan 19, 2026Yuhiro Ono, Tomohiro Harada, Yukiya MiuraEvolutionary Optimization Methods

  3. AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation

    Jun 30, 2026Xinyuan Song, Zekun Cai, Liang ZhaoCode Generation