Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing rules. Because these contents jointly shape later execution, a harness update can disrupt previously reliable behavior even when the model is frozen. This raises a new question: how can an agent continually improve its state outside the model while retaining behavior acquired earlier? We formulate Harness Continual Learning (HCL), a new continual learning paradigm in which the harness evolves around a frozen foundation model, and define the resulting loss of earlier behavior as harness-level forgetting. We instantiate HCL with four execution-facing components: the Task Interface, Experience Memory, Capability Map, and Adaptive Router. We further introduce guarded harness evolution to separate update generation from state commitment. A Continual Optimizer proposes candidate harnesses from post-execution feedback, and a Continual Evaluator commits the resulting candidate harness only after checking current improvement, historical retention, and validity. Experiments on textual reasoning, multimodal perception, and open-world interaction demonstrate capability accumulation and failure recovery, with relative gains exceeding 10% over corresponding baselines in multiple settings. Component ablations assess the contribution of each harness component, while controlled retention sweeps reveal measurable harness-level forgetting and show that the stability--plasticity trade-off can be explicitly adjusted.
Figures & tables
Figure 1: The shift in the object of continual learning. Model-centric methods update model parameters θ over sequential experience. HCL instead updates harness state around a frozen foundation model. In both settings, adaptation can improve performance on new capabilities, but may also degrade capabilities acquired earlier.
Figure 2: Overview of HCL. The harness Hn supports execution from interaction un to outcome yn . After execution, the Continual Optimizer proposes candidate harnesses, which the Continual Evaluator accepts or rejects based on current improvement, historical retention, and validity.
Method
Pick
Look
Clean
Heat
Cool
Two-object
Final Avg. ↑
Avg. Fgt. (%) ↓
Static Harness
95.80
66.70
25.80
26.10
9.50
58.80
47.12
–
RAG Baseline
95.80
83.30
41.90
39.10
14.30
58.80
55.56
1.74
MemP ( Fang et al., 2026 )
95.80
83.30
48.40
34.80
9.50
47.10
53.15
5.18
MemRL ( Zhang et al., 2026c )
95.83
88.89
41.94
39.13
57.14
41.18
60.68
1.76
Stability-HCL (Ours)
100.00
83.30
51.60
30.40
28.60
76.50
61.74
2.64
Plasticity-HCL (Ours)
100.00
77.80
41.90
39.10
19.00
100.00
62.98
10.94
Table 1: Final task-wise performance on ALFWorld. The best and second-best results in each metric column are marked in bold and underlined , respectively.
Figure 3: Minecraft curriculum progression and execution efficiency. (a) Plasticity-HCL completes all 50 tasks, while the Static Harness plateaus at 15. (b) Plasticity-HCL uses 83 cumulative environment actions, versus 88 for MemRL and 91 for MemP; lower is more efficient.
Textual reasoning: MuSiQue → ProofWriter → GSM8K → HotpotQA
Method
MuSiQue
ProofWriter
GSM8K
HotpotQA
Final Avg. ↑
Avg. Fgt. (%) ↓
DeepSeek-V4.1-Flash Zero-shot
48.00
79.20
65.20
62.20
63.65
–
MemP ( Fang et al., 2026 )
51.40
78.80
70.80
60.80
65.45
0.47
MemRL ( Zhang et al., 2026c )
53.80
80.00
77.40
64.20
68.85
0.07
Plasticity-HCL (Ours)
58.60
85.60
96.00
64.20
76.10
0.40
Stability-HCL (Ours)
53.60
82.80
82.60
63.60
70.65
0.27
Table 2: Final task-wise performance on the textual and multimodal streams.
b
MuSiQue
ProofWriter
GSM8K
HotpotQA
Final Avg. ↑
Avg. Fgt. (%) ↓
0
27.83
73.33
84.33
59.50
61.25
0.39
1
24.83
77.50
92.33
59.17
63.46
1.22
3
26.83
79.83
83.00
58.50
62.04
2.00
∞
28.33
71.00
82.00
59.17
60.13
3.45
Table 3: Effect of historical-loss tolerance b on final performance and forgetting.
Figure 4: Effect of stronger-model placement in HCL. Base-HCL is fully driven by Qwen3-4B. Opt-HCL uses Qwen3.6-27B only for candidate harness generation while keeping Qwen3-4B for online execution. Online-HCL uses Qwen3.6-27B for online HCL operations. (a) Mean online processing time per matched sample with one standard deviation. (b) Final four-task average accuracy.
Component
I
M
C
R
Final Avg. ↑
Avg. Fgt. (%) ↓
Zero-shot
–
–
–
–
34.84
–
w/o Interface update
×
✓
✓
✓
62.37
0.11
w/o Memory update
✓
×
✓
✓
62.28
0.83
w/o Capability update
✓
✓
×
✓
63.12
0.06
w/o Router update
✓
✓
✓
×
62.77
0.14
Plasticity-HCL
✓
✓
✓
✓
63.41
0.45
Table 4: Component ablation on the controlled multimodal stream. A component is kept available during execution but its persistent updates are frozen.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
MuSiQue
ProofWriter
GSM8K
HotpotQA
Final Avg. ↑
Avg. Fgt. (%) ↓
Qwen3.6-27B Zero-shot
57.00
83.80
44.80
62.40
62.00
–
MemP ( Fang et al., 2026 )
56.40
74.00
54.00
65.80
62.55
0.20
MemRL ( Zhang et al., 2026c )
55.80
75.20
52.20
65.60
62.20
0.00
Plasticity-HCL (Ours)
56.00
79.80
66.00
62.80
66.15
7.87
Stability-HCL (Ours)
55.80
77.00
51.00
64.20
62.00
1.27
Appendix
Table 5: Textual-reasoning results with frozen Qwen3.6-27B. The best and second-best entries in each column are marked in bold and underlined , respectively.
Figure 5: Stage-wise forgetting under different historical-loss tolerances. (a) Textual reasoning comparing Stability-HCL and Plasticity-HCL. Forgetting values are reported in percent. (b) Multimodal perception comparing Stability-HCL and Plasticity-HCL. Forgetting values are reported in percent.
Figure 6: Capability-aware HCL with an external arithmetic calculator. (a) Final accuracy compared with the earlier Stability-HCL run. (b) Corresponding accuracy gains; GSM8K shows the largest improvement ( +37.20 pp), with 254/500 examples routed to the calculator. The comparison is exploratory because the runs also differ in prompts, retries, and memory trajectories.
Method
I
M
C
R
Fixed contents
Zero-shot
–
–
–
–
No structured HCL harness or persistent updates.
Full HCL
✓
✓
✓
✓
None.
w/o Interface update
×
✓
✓
✓
Prompts, templates, parsing, and normalization rules.
w/o Memory update
✓
×
✓
✓
Raw and Abstract Memory entries.
w/o Capability update
✓
✓
×
✓
Reusable skills.
w/o Router update
✓
✓
✓
×
Routing prompts, selection criteria, and workflow templates.
Appendix
Table 6: Update scope of the component-ablation variants. A ✓ permits persistent updates, while × keeps the corresponding component fixed.
Method
Detection
Caption
Grounding
VQAv2
Final Avg. ↑
Avg. Fgt. (%) ↓
Zero-shot
35.11
22.98
0.00
81.27
34.84
–
w/o Interface update
53.45
33.56
87.80
74.67
62.37
0.11
w/o Memory update
55.50
28.95
88.00
76.67
62.28
0.83
w/o Capability update
55.11
34.16
86.40
76.80
63.12
0.06
w/o Router update
53.59
36.68
87.40
73.40
62.77
0.14
Full HCL
53.07
36.09
87.60
76.87
63.41
0.45
Appendix
Table 7: Full component-ablation results on the controlled multimodal stream. “Committed” counts candidate updates entering the persistent harness.