cs.AISep 27, 2026

Vestrum: Improving Agent Harnesses by Adapting Their Verification, Structure and Memory

Authors: Jayant Parashar, Eugene F. Douglass, William C. Bastian, Suchendra M. Bhandarkar

Organizations: University of Georgia · Emory University

Abstract

An agent harness controls how a language model accesses information, uses tools, preserves memory, and checks its work. Improving this software is costly when each evaluation requires a long interaction with an environment. We introduce Vestrum, a framework that turns failures in execution traces into scoped harness changes without training the task model. Its organizing overhypothesis is that tasks of a shared kind may exhibit recurring failures whose remedies transfer within that kind. Vestrum expresses failures as recognizable classes, proposes changes across verification, retrieval, decomposition, and knowledge synthesis, and screens their scope before evaluating them as a bundle. A persistent lessons file informs subsequent proposals. Across five settings and two baseline harnesses, the frozen harnesses improve held-out performance: UltraHorizon rises from 47.6 to 59.8 over GAM, Terminal-Bench 4 Hard from 63.7% to 70.3% of checks passed over Claude Code on eight held-out tasks at 1.03x test cost, and cell-type annotation agreement from 67.5% to 77.8% on held-out sections of one slide, alongside gains on LoCoMo and AMA-Bench. Across our searches, verification grounded in evidence helped both intermediate steps and final answers, at lower cost at intermediate steps, while critics asked to rebuild finished answers broke more than they repaired. On the three memory benchmarks, Vestrum also scores above the evaluated GEPA configurations in every paired evaluation.

Explore similar work

CardsList
  1. Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

    May 27, 2026Yilun Yao, Xinyu Tan, Chao-Hsuan Liu +9Agent HarnessAgentic Benchmarks

  2. VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

    Oct 1, 2026Caiqi Zhang, Rujun Han, Zifeng Wang +4Long-Horizon AgentsLong-Horizon Task Planning

  3. A2EA^2E : An End-to-End Agent Auditing Engine

    Aug 7, 2026Haoning Wang, Mingxun Zhang, Chenyue Yu +4Agent HarnessAgentic Evaluations