MetaBench-Harness: Unlocking End-to-End Optimization of Benchmark Harnesses
Organizations: National Taiwan University · NTU AI-Core · Queen Mary University of London
Abstract
Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving beyond the evolution of isolated tasks, we propose to optimize the benchmark generation workflow itself end to end with MetaBench-Harness, a dual-loop search framework. Specifically, the inner loop utilizes a benchmark harness to generate a new benchmark in each round, while the outer meta-harness orchestration layer iteratively refines and searches over harness implementations based on historical evolution trajectories. By applying MetaBench-Harness to the competitive programming CodeContests and Olympiad mathematics AIME-2024 datasets, we demonstrate that the evolved benchmarks are challenging and discriminative for frontier models. Trajectory and quality analyses verify that MetaBench-Harness enables multi-dimensional evolution, steadily improving evolution reasonableness, benchmark competency, and evaluator robustness across successive rounds. Furthermore, case studies reveal its effective utilization of diverse difficulty levers to reframe problems and elevate required capabilities. Ultimately, this work provides a solution to the pressing challenge of benchmark saturation.
Figures & tables
| Evolution Methods | Note | DeepSeek | Kimi | Gemini | Qwen-32B | Qwen-7B |
| Baseline | AIME-2024 | 0.900 | 0.900 | 0.833 | 0.700 | 0.533 |
| TRACE | Reported | N/A | N/A | N/A | 0.533 | 0.393 |
| AutoEvoEval | Reproduced | 0.700 | 0.633 | 0.500 | 0.400 | 0.367 |
| Naïve Prompting | Ablation | 1.000 | 1.000 | 1.000 | 0.875 | 0.750 |
| Fixed Harness | Ablation | 0.817 | 0.767 | 0.633 | 0.600 | 0.133 |
| 0.322 | 0.544 | 0.144 | 0.089 | 0.044 |
| Model | Type | Loop | MB-CodeContests | MB-AIME | ||||
| Base. | Early | Late | Base. | Early | Late | |||
| Claude-Sonnet-5 | AR | ✓ | 0.835 | 0.948 | 0.839 | 0.733 | 0.750 | 0.633 |
| GPT-5.4 | NR | ✓ | 0.586 | 0.591 | 0.293 | 0.067 | 0.125 | 0.100 |
| DeepSeek-V3.1 | NR | 0.736 | 0.845 | 0.396 | 0.667 | 0.125 | 0.133 | |
| Kimi-K2.5 | NR | 0.694 | 0.834 | 0.739 | 0.833 | 0.333 | 0.167 | |
| Gemini-2.5-Flash | NR | 0.458 | 0.484 | 0.278 | 0.767 | 0.167 | 0.100 | |
| Model | Type | MB-CodeContests | MB-AIME | ||
| Tokens | Tokens | ||||
| DeepSeek-V3.1 | R | ||||
| Kimi-K2.5 | R | ||||
| Gemini-2.5-Flash | R | ||||
| MB-CodeContests | MB-AIME | |
| Meta-Harness Orchestrated Proposals | ||
| Optimization rounds | 10 | 5 |
| Proposals applied | 20/23 (87%) | 10/12 (83%) |
| Edits Applied by the Harness | ||
| Measurement | 11 (39%) | 2 (20%) |
| Generation | 9 (32%) | 6 (60%) |
| MB-CodeContests | MB-AIME | |
| Meta-Harness Orchestrated Proposals | ||
| Optimization rounds | 10 | 5 |
| Proposals applied | 20/23 (87%) | 10/12 (83%) |
| Edits Applied by the Harness | ||
| Measurement | 11 (39%) | 2 (20%) |
| Generation | 9 (32%) | 6 (60%) |
| #. Ori. | #. Evo. | |
| MB-CodeContests | ||
| Total Questions | 165 | 100 |
| Solving Techniques | 5.02 | 5.49 |
| Capability | 3.55 | 3.90 |
| MB-AIME | ||
| Total Questions | 30 | 30 |
| Metric | MB-CC ∗ | MB-AIME | Human |
| Reasonableness | 95% / 91% | 90% / 83% | 85% |
| Competency | 78% / 61% | 80% / 78% | 68% |
| Robustness | 90% / 62% | 93% / 69% | 85% |
| Overall | 70% / 44% | 70% / 49% | 80% |
| ∗ MB-CC denotes MB-CodeContests. | |||
| Part I: Case Study of MB-CodeContests Evolution | ||
| Seed (Codeforces 1582D) | Round 5 (Evolved Emergence) | Round 9 (Evolved Deepening) |
| Task Formulation | ||
| Constructive Math : Given an array ( ), construct such that and . | Single Gap Audit : Exploit a planted gap ( range(n-1) ) to trick a Python checker with an invalid pair. | Compound Audit : Trace complex logic to exploit a vulnerability gated by 3 conditions ( , duplicate , odd ). |
| Evolution Summary | ||
| Capability: 3 Technique: 3 Lever: Baseline specification Baseline: The root foundation originating from a standard coding task. | Capability: 3 Technique: 5 Key Lever: Meta level shift, Output contract tightening Moderate: Shifts to basic code auditing and static inspection. | Capability: 4 Technique: 6 Key Lever: Compound conjunction, Scaffolding removal High: Mandates multi-technique fusion without hints. |
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
| Part I: Case Study of MB-CodeContests Evolution | ||
| Seed (Codeforces 1582D) | Round 5 (Evolved Emergence) | Round 9 (Evolved Deepening) |
| Task Formulation | ||
| Constructive Math : Given an array ( ), construct such that and . | Single Gap Audit : Exploit a planted gap ( range(n-1) ) to trick a Python checker with an invalid pair. | Compound Audit : Trace complex logic to exploit a vulnerability gated by 3 conditions ( , duplicate , odd ). |
| Capabilities and Solving Techniques | ||
| Capabilities: Construction search, Domain reasoning, Precision scale control Solving Techniques: Constructive algorithm design, Exact integer arithmetic, Magnitude bound control | Capabilities: Adversarial analysis, Construction search, Output contract Solving Techniques: Code reading static analysis, Checker vulnerability exploitation, Counterexample construction, Program execution simulation, Protocol compliance | Capabilities: Adversarial analysis, Construction search, Precision scale control, Output contract Solving Techniques: Code reading static analysis, Checker vulnerability exploitation, Counterexample construction, Long horizon arithmetic stamina, Casework exhaustiveness, Protocol compliance |
| Applied Difficulty Levers | ||