Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean partial-credit score of 75.19% on AutomationBench, exceeding weight-only RL by 6.43 and 9.09 percentage points and ungated alternation by 4.59 and 6.95 points, respectively. Across 44 harness proposals, 17 reduce validation performance at the updated checkpoint and are rejected before subsequent RL training, highlighting the importance of validation gating.
Figures & tables
Figure 1: Conceptual schedules for model–harness co-optimization. Harness-only and weight-only update one component. SIA-style and WHALE-style illustrate stagewise and alternating schedules, respectively; VACE validates proposed harness updates before the next training round. The figure uses θ,h for the weights W and harness H in our notation. The paths and contours are schematic, and the star marks a conceptual joint optimum.
1:
for t=0,…,T−1 do
2:
(Wt+1,Tt)←M(Wt;Ht,Dtrain)
3:
Ht′←S(Ht,Tt)
4:
bt←RV(Wt+1,Ht)
ct←RV(Wt+1,Ht′)
5:
if ct>bt then
6:
Ht+1←Ht′
Algorithm 1 VACE: alternating updates with validation-gated harness acceptance
Split
HR
Marketing
Finance
Total
Train
26
27
27
80
Validation
15
22
21
58
Test
24
23
24
71
Total
65
72
72
209
Table 1: AutomationBench task counts by domain and split used in our experiments.
Method
OfficeQA
AutomationBench
HR
Marketing
Finance
Overall
Static agent
32.11
43.86
52.57
54.44
50.26
Harness-only
35.28
55.03
64.53
65.42
61.62
RL-only
38.83
63.98
67.55
66.85
66.10
SIA-style [ Hebbar et al., 2026 ]
40.67
56.36
68.83
71.01
65.35
WHALE-style [ Kim et al., 2026 ]
40.67
61.50
72.30
71.11
68.25
Table 2: Mean test scores (%) from three evaluation runs of each selected model–harness pair: accuracy on OfficeQA and mean partial credit on AutomationBench. Bold marks the highest listed mean in each column.
Table 4: Same-checkpoint examples from the phase-level trajectories: the first accepted and first rejected proposal in each dataset. Scores are percentages; changes are percentage points.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Method
RL cap
H cap
Harness gate
Static agent
0
0
—
Harness-only
0
BH
Yes
RL-only
BW
0
—
SIA-style
BW
BH
Yes
WHALE-style
BW
BH
No
VACE
BW
BH
Yes
Appendix
Table 5: Optimization settings. BW and BH denote the planned RL and harness-proposal caps.