cs.LGSep 29, 2026

VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses

Authors: Jiexing Qi, Yu He, Jun Liu, Qichen Huang, Shaohua Hu, Zhan Dang, Guohua Chen, Rui Yang, +4 more

Organizations: ICT AI Competence Center, Huawei Technologies Co., Ltd., Shanghai, China · Shanghai Jiao Tong University, Shanghai, China

Abstract

Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean partial-credit score of 75.19% on AutomationBench, exceeding weight-only RL by 6.43 and 9.09 percentage points and ungated alternation by 4.59 and 6.95 points, respectively. Across 44 harness proposals, 17 reduce validation performance at the updated checkpoint and are rejected before subsequent RL training, highlighting the importance of validation gating.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. WHALE: A Simple Recipe for Joint Harness-Weight Optimization

    Aug 31, 2026Haechan Kim, Yoonho Lee, Gisang Lee +2Agent HarnessAgentic Optimization

  2. EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

    Aug 6, 2026Zishan Xu, Zhiyuan Yao, Yuxin Chen +9Agentic Reinforcement LearningLarge Language Model Agents