Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoEvolution, which alternates agentic reinforcement learning with trajectory-driven harness refinement. After each RL stage, VACE reuses the collected trajectories to propose a harness revision and evaluates the incumbent and candidate with the updated model held fixed. The candidate guides subsequent training only if it improves validation performance. With Qwen3.5-9B, VACE achieves 45.26% test accuracy on OfficeQA and a mean partial-credit score of 75.19% on AutomationBench, exceeding weight-only RL by 6.43 and 9.09 percentage points and ungated alternation by 4.59 and 6.95 points, respectively. Across 44 harness proposals, 17 reduce validation performance at the updated checkpoint and are rejected before subsequent RL training, highlighting the importance of validation gating.
Figures & tables
Figure 1: Conceptual schedules for model–harness co-optimization. Harness-only and weight-only update one component. SIA-style and WHALE-style illustrate stagewise and alternating schedules, respectively; VACE validates proposed harness updates before the next training round. The figure uses θ,h for the weights W and harness H in our notation. The paths and contours are schematic, and the star marks a conceptual joint optimum.
1:
for t=0,…,T−1 do
2:
(Wt+1,Tt)←M(Wt;Ht,Dtrain)
3:
Ht′←S(Ht,Tt)
4:
bt←RV(Wt+1,Ht)
ct←RV(Wt+1,Ht′)
5:
if ct>bt then
6:
Ht+1←Ht′
Algorithm 1 VACE: alternating updates with validation-gated harness acceptance
Split
HR
Marketing
Finance
Total
Train
26
27
27
80
Validation
15
22
21
58
Test
24
23
24
71
Total
65
72
72
209
Table 1: AutomationBench task counts by domain and split used in our experiments.
Method
OfficeQA
AutomationBench
HR
Marketing
Finance
Overall
Static agent
32.11
43.86
52.57
54.44
50.26
Harness-only
35.28
55.03
64.53
65.42
61.62
RL-only
38.83
63.98
67.55
66.85
66.10
SIA-style [ Hebbar et al., 2026 ]
40.67
56.36
68.83
71.01
65.35
WHALE-style [ Kim et al., 2026 ]
40.67
61.50
72.30
71.11
68.25
Table 2: Mean test scores (%) from three evaluation runs of each selected model–harness pair: accuracy on OfficeQA and mean partial credit on AutomationBench. Bold marks the highest listed mean in each column.
Table 4: Same-checkpoint examples from the phase-level trajectories: the first accepted and first rejected proposal in each dataset. Scores are percentages; changes are percentage points.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Method
RL cap
H cap
Harness gate
Static agent
0
0
—
Harness-only
0
BH
Yes
RL-only
BW
0
—
SIA-style
BW
BH
Yes
WHALE-style
BW
BH
No
VACE
BW
BH
Yes
Appendix
Table 5: Optimization settings. BW and BH denote the planned RL and harness-proposal caps.
Agent performance depends jointly on the model parameters and the executable harness code that manages context and control flow. Optimizing either component in isolation can leave the system bottlenecked by its frozen counterpart: weight updates can change which harness is effective, while harness updates can change which model capabilities are exposed. Existing joint-adaptation methods optimize weights and textual prompts but leave the broader harness fixed. We propose Weight-Harness Alternating LEarning (WHALE), a simple recipe that alternates two phases: updating the model under the current harness, then searching for a better harness under the updated model. We instantiate these two phases with online rejection-sampling fine-tuning and Meta-Harness, respectively. When to switch is a key design choice: to separate real improvements from noise without over-optimizing against a changing counterpart, WHALE uses either fixed phase durations or an adaptive patience rule over training signals. Using Qwen3.5-2B/4B agents across three domains (search question answering, mathematical reasoning, and chess puzzles), WHALE outperforms weight-only, harness-only, and Fast-Slow Training by 4.15-24.38 percentage points in best mean@8 accuracy. Either component can be the bottleneck: harness search matches peak weight-only accuracy with far fewer rollouts in SearchQA, but improves math accuracy only after a weight update. Small interleaved updates also outperform stagewise weight-then-harness optimization in accuracy and rollout cost. The code is available at https://github.com/krafton-ai/WHALE.
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. The policy alternates between acting and rehearsal: it first generates a tool call, then plays the role of the environment to produce the response induced by that action, and conditions subsequent decisions on the rehearsed response. Both roles are jointly optimized end-to-end using task-success rewards. Through world rehearsal, the policy internalizes the relationship between actions and their environment responses in its parameters, yielding an agent world model that directly supports decision making. Across BFCL-v4, tau^2-Bench, VitaBench, and FinMCP-Bench, EnvACE achieves strong and transferable performance, outperforming environment-scaling baselines in the overall evaluation. Controlled studies further show that world rehearsal consistently improves policy learning across model scales. At test time, the internalized world model enables private rehearsal before committed execution, yielding further gains under a moderate rehearsal budget without additional external interaction. Our findings establish world rehearsal as a new path toward scaling LLM agent training beyond the constraints of external environments. Our code is publicly available at https://github.com/Within-yao/EnvACE.
Zishan Xu, Zhiyuan Yao, Yuxin Chen +9
Shanghai Jiao Tong University · Zhejiang University · National University of Singapore +4
Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other components of the agent pipeline. Self-evolution offers a promising alternative, but most existing frameworks assume access to frontier models that can reliably diagnose failures, propose revisions, and judge their own updates. We study whether frozen small language models (SLMs) can serve as effective self-evolving agents under resource constraints. We propose PACE (Prompt And Control Logic Evolution), a two-timescale framework that coordinates low-risk prompt refinement with higher-risk control-logic updates. PACE evolves prompts under fixed control logic until prompt-level gains saturate, then considers constrained control-logic updates that are accepted through held-out validation. Across three frozen SLM backbones ranging from 4B to 14B parameters and four controlled benchmarks, PACE achieves the best performance on all 12 backbone--benchmark combinations, improving over vanilla SLM agents by up to +9.2% relative improvement and over the stronger single-mode evolution baseline by up to +5.4% relative improvement. A tau-bench case study further shows that PACE improves multi-turn tool-use success over vanilla and prompt-only evolution. These results suggest that reliable SLM agent self-evolution is possible without updating model weights or relying on frontier-model teachers, and that the key benefit is not any single final solver pattern but autonomous, validated discovery of task-appropriate inference strategies.