cs.LGOct 4, 2026

What Does a Harness Repair? A Preregistered Study of Visibility, Baseline Adequacy and Evaluation Defects

Authors: Bowen Xu, Boyu Chen

Organizations: Department of Biology, Stanford University, Stanford, CA, USA · Institute of Health Informatics, University College London, London, UK

Abstract

Harness search keeps a change to the prompts, reasoning switches, token budgets or parsers around a frozen model if the change raises a score. Such a gain can come from answers the parser could not read before, a weak comparison, or a defect in the evaluation. We preregistered a study of where these gains come from, with three small models, three benchmarks, replication and test partitions, a GEPA search arm and six evaluation defects injected one at a time, and we report all 47 primary endpoints. Turning thinking off raised accuracy over a capped thinking setting in 5 of 9 model-benchmark cells, and in each the gain came mostly from questions where the capped setting gave no readable answer. The thinking-off setting was not meaningfully worse than a rescue configuration or four GEPA-selected harnesses in 11 of 13 comparisons, and lost to the rescue on GSM8K for two models. GEPA repaired its broken starting points, but none of its selected harnesses was more accurate than the thinking-off setting. A thinking budget in the serving engine, which also allows a longer answer, lowered truncation and raised the parse rate in 6 of 9 cells. In 6 of 15 evaluable defect-model pairs, replication through the same pipeline reproduced the defect's distortion instead of revealing it. On the LongevityBench multiple-choice tasks, only the longevity-tuned model beat the strongest constant-label baseline.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Finding the Right Fit: Model-Harness Interactions across Agent Tasks

    Oct 1, 2026Yixuan Li, Yiyun Zhou, Yao Long Teng +6Tool-Augmented Language Model AgentsLLM Agent Evaluation

  2. LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

    Sep 29, 2026Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen +1Long-Context Language Model EvaluationLong-Context Language Model Inference