Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing rules. Because these contents jointly shape later execution, a harness update can disrupt previously reliable behavior even when the model is frozen. This raises a new question: how can an agent continually improve its state outside the model while retaining behavior acquired earlier? We formulate Harness Continual Learning (HCL), a new continual learning paradigm in which the harness evolves around a frozen foundation model, and define the resulting loss of earlier behavior as harness-level forgetting. We instantiate HCL with four execution-facing components: the Task Interface, Experience Memory, Capability Map, and Adaptive Router. We further introduce guarded harness evolution to separate update generation from state commitment. A Continual Optimizer proposes candidate harnesses from post-execution feedback, and a Continual Evaluator commits the resulting candidate harness only after checking current improvement, historical retention, and validity. Experiments on textual reasoning, multimodal perception, and open-world interaction demonstrate capability accumulation and failure recovery, with relative gains exceeding 10% over corresponding baselines in multiple settings. Component ablations assess the contribution of each harness component, while controlled retention sweeps reveal measurable harness-level forgetting and show that the stability--plasticity trade-off can be explicitly adjusted.
Figures & tables
Figure 1: The shift in the object of continual learning. Model-centric methods update model parameters θ over sequential experience. HCL instead updates harness state around a frozen foundation model. In both settings, adaptation can improve performance on new capabilities, but may also degrade capabilities acquired earlier.
Figure 2: Overview of HCL. The harness Hn supports execution from interaction un to outcome yn . After execution, the Continual Optimizer proposes candidate harnesses, which the Continual Evaluator accepts or rejects based on current improvement, historical retention, and validity.
Method
Pick
Look
Clean
Heat
Cool
Two-object
Final Avg. ↑
Avg. Fgt. (%) ↓
Static Harness
95.80
66.70
25.80
26.10
9.50
58.80
47.12
–
RAG Baseline
95.80
83.30
41.90
39.10
14.30
58.80
55.56
1.74
MemP ( Fang et al., 2026 )
95.80
83.30
48.40
34.80
9.50
47.10
53.15
5.18
MemRL ( Zhang et al., 2026c )
95.83
88.89
41.94
39.13
57.14
41.18
60.68
1.76
Stability-HCL (Ours)
100.00
83.30
51.60
30.40
28.60
76.50
61.74
2.64
Plasticity-HCL (Ours)
100.00
77.80
41.90
39.10
19.00
100.00
62.98
10.94
Table 1: Final task-wise performance on ALFWorld. The best and second-best results in each metric column are marked in bold and underlined , respectively.
Figure 3: Minecraft curriculum progression and execution efficiency. (a) Plasticity-HCL completes all 50 tasks, while the Static Harness plateaus at 15. (b) Plasticity-HCL uses 83 cumulative environment actions, versus 88 for MemRL and 91 for MemP; lower is more efficient.
Textual reasoning: MuSiQue → ProofWriter → GSM8K → HotpotQA
Method
MuSiQue
ProofWriter
GSM8K
HotpotQA
Final Avg. ↑
Avg. Fgt. (%) ↓
DeepSeek-V4.1-Flash Zero-shot
48.00
79.20
65.20
62.20
63.65
–
MemP ( Fang et al., 2026 )
51.40
78.80
70.80
60.80
65.45
0.47
MemRL ( Zhang et al., 2026c )
53.80
80.00
77.40
64.20
68.85
0.07
Plasticity-HCL (Ours)
58.60
85.60
96.00
64.20
76.10
0.40
Stability-HCL (Ours)
53.60
82.80
82.60
63.60
70.65
0.27
Table 2: Final task-wise performance on the textual and multimodal streams.
b
MuSiQue
ProofWriter
GSM8K
HotpotQA
Final Avg. ↑
Avg. Fgt. (%) ↓
0
27.83
73.33
84.33
59.50
61.25
0.39
1
24.83
77.50
92.33
59.17
63.46
1.22
3
26.83
79.83
83.00
58.50
62.04
2.00
∞
28.33
71.00
82.00
59.17
60.13
3.45
Table 3: Effect of historical-loss tolerance b on final performance and forgetting.
Figure 4: Effect of stronger-model placement in HCL. Base-HCL is fully driven by Qwen3-4B. Opt-HCL uses Qwen3.6-27B only for candidate harness generation while keeping Qwen3-4B for online execution. Online-HCL uses Qwen3.6-27B for online HCL operations. (a) Mean online processing time per matched sample with one standard deviation. (b) Final four-task average accuracy.
Component
I
M
C
R
Final Avg. ↑
Avg. Fgt. (%) ↓
Zero-shot
–
–
–
–
34.84
–
w/o Interface update
×
✓
✓
✓
62.37
0.11
w/o Memory update
✓
×
✓
✓
62.28
0.83
w/o Capability update
✓
✓
×
✓
63.12
0.06
w/o Router update
✓
✓
✓
×
62.77
0.14
Plasticity-HCL
✓
✓
✓
✓
63.41
0.45
Table 4: Component ablation on the controlled multimodal stream. A component is kept available during execution but its persistent updates are frozen.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
MuSiQue
ProofWriter
GSM8K
HotpotQA
Final Avg. ↑
Avg. Fgt. (%) ↓
Qwen3.6-27B Zero-shot
57.00
83.80
44.80
62.40
62.00
–
MemP ( Fang et al., 2026 )
56.40
74.00
54.00
65.80
62.55
0.20
MemRL ( Zhang et al., 2026c )
55.80
75.20
52.20
65.60
62.20
0.00
Plasticity-HCL (Ours)
56.00
79.80
66.00
62.80
66.15
7.87
Stability-HCL (Ours)
55.80
77.00
51.00
64.20
62.00
1.27
Appendix
Table 5: Textual-reasoning results with frozen Qwen3.6-27B. The best and second-best entries in each column are marked in bold and underlined , respectively.
Figure 5: Stage-wise forgetting under different historical-loss tolerances. (a) Textual reasoning comparing Stability-HCL and Plasticity-HCL. Forgetting values are reported in percent. (b) Multimodal perception comparing Stability-HCL and Plasticity-HCL. Forgetting values are reported in percent.
Figure 6: Capability-aware HCL with an external arithmetic calculator. (a) Final accuracy compared with the earlier Stability-HCL run. (b) Corresponding accuracy gains; GSM8K shows the largest improvement ( +37.20 pp), with 254/500 examples routed to the calculator. The comparison is exploratory because the runs also differ in prompts, retries, and memory trajectories.
Method
I
M
C
R
Fixed contents
Zero-shot
–
–
–
–
No structured HCL harness or persistent updates.
Full HCL
✓
✓
✓
✓
None.
w/o Interface update
×
✓
✓
✓
Prompts, templates, parsing, and normalization rules.
w/o Memory update
✓
×
✓
✓
Raw and Abstract Memory entries.
w/o Capability update
✓
✓
×
✓
Reusable skills.
w/o Router update
✓
✓
✓
×
Routing prompts, selection criteria, and workflow templates.
Appendix
Table 6: Update scope of the component-ablation variants. A ✓ permits persistent updates, while × keeps the corresponding component fixed.
Method
Detection
Caption
Grounding
VQAv2
Final Avg. ↑
Avg. Fgt. (%) ↓
Zero-shot
35.11
22.98
0.00
81.27
34.84
–
w/o Interface update
53.45
33.56
87.80
74.67
62.37
0.11
w/o Memory update
55.50
28.95
88.00
76.67
62.28
0.83
w/o Capability update
55.11
34.16
86.40
76.80
63.12
0.06
w/o Router update
53.59
36.68
87.40
73.40
62.77
0.14
Full HCL
53.07
36.09
87.60
76.87
63.41
0.45
Appendix
Table 7: Full component-ablation results on the controlled multimodal stream. “Committed” counts candidate updates entering the persistent harness.
A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, without performing any parameter-space update. Experiments on reasoning and multi-hop question answering show that harness learning improves revision quality and that the ability to adapt at test time transfers to unseen tasks. Policies trained on individual revisions can continue improving harnesses over multiple rounds, while the benefits of training on revision sequences vary across settings. These findings suggest a path towards continually learning agents that turn accumulated experience into generalizable improvements.
Can language agents continually learn from experience, turning earlier interactions into reusable capabilities? AhaBench evaluates this ability through exploration after solved hidden-state puzzles, computational transfer after mathematical teaching, and sustained business operation under delayed feedback. The benchmark is agnostic to how an agent learns; the evaluated agents use fixed model weights. Curriculum profiles, teaching contrasts, and daily trajectories reveal a common challenge: using explicit guidance is more reliable than generalizing beyond it or sustaining useful behavior. Across the Puzzle panel, the advantage over matched cold targets is 36.0-53.5 points greater with trace support than at the trace-free endpoint; Qwen 3.6 Plus nevertheless retains a +12.57-point post-curriculum gain. In Euler, worked procedures yield 80.0-100.0% held-out accuracy across models, while question-plus-answer teaching yields 0.0-73.9%. Vending trajectories separate sustained profit, late recovery, and incomplete operation: Doubao Seed 2.0 Pro finishes nominal operation at +495 but averages -10 over the year. Together, these results make continual learning an operational target: experience should yield capabilities that remain effective as guidance, inputs, and business states change. We release tasks, validators, a simulator, records, and analyses for developing agents that turn useful insights into lasting abilities.
Zerui Cheng, Jiawei Xu, Huacan Chai +3
Princeton University · Tencent Hy · Tsinghua University +2
Coding harnesses such as Claude Code and OpenHands wrap foundation models with tools, memory, and planning, but no equivalent exists for embodied agents' long-horizon partial-observability decision-making. We first report our Gemini Plays Pokemon (GPP) experiments. With iterative human-in-the-loop harness refinement, GPP became the first AI system to complete Pokemon Blue, Yellow Legacy on hard mode, and Crystal without a lost battle. In the hardest stages, the agent itself began iterating on its strategy through long-context memory, surfacing emergent self-improvement signals alongside human-in-the-loop refinement. Continual Harness removes the human fully from this loop: a reset-free self-improving harness for embodied agents that formalizes and automates what we observed. Starting from only a minimal environment interface, the agent alternates between acting and refining its own prompt, sub-agents, skills, and memory, drawing on any past trajectory data. Prompt-optimization methods require episode resets; Continual Harness adapts online within a single run. On Pokemon Red and Emerald across frontier models, Continual Harness starting from scratch substantially reduces button-press cost relative to the minimalist baseline and recovers a majority of the gap to a hand-engineered expert harness, with capability-dependent gains, despite starting from the same raw interface with no curated knowledge, no hand-crafted tools, and no domain scaffolding. We then close the loop with the model itself: an online process-reward co-learning loop, in which an open-source agent's rollouts through the refining harness are relabeled by a frontier teacher and used to update the model, drives sustained in-game milestone progress on Pokemon Red without resetting the environment between training iterations.
Seth Karten, Joel Zhang, Tersoo Upaa +5
Princeton University · ARISE Foundation · Google DeepMind