A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT
Figures & tables
Figure 1 : Imitation and reinforcement lose executable behavior. Online per-episode updates (Sec. 3.2 ) of Qwen3.5-4B on the ALFWorld seen stream, with std shading: (a) first-response-token entropy, (b) KL to the initial policy, (c) valid-action rate, and (d) generated tokens per turn.
Figure 2 : Online Agentic Test-Time Training of ASCENT. The student makes a single attempt per task. A verified trajectory becomes privileged information for the frozen teacher, whose soft targets update the LoRA fast weights for the next task. Charts on the right show ALFWorld seen-stream means of 4B and 9B runs (per-scale in Fig. 5 (a) and Table 1 ). Episode time includes the update.
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg. SR
Turns
Time GPU·h / s
Qwen3.5-4B
Base
91.4
46.2
40.7
6.2
8.0
54.2
46.4
35.1
97
Offline
MemP
94.3
53.8
77.8
0.0
36.0
66.7
61.4 ± 6.5
28.6
5.2 / 61
ACE
85.7
46.2
70.4
6.2
60.0
58.3
60.7 ± 9.2
31.1
6.3 / 81
EvoSkill
93.3
17.9
35.8
6.2
32.0
66.7
49.8 ± 4.9
34.4
5.9 / 123
Table 1: ALFWorld seen-stream results. Success rate (SR, %) by task type and over the full stream, mean turns per episode, and runtime, given as training GPU-hours / evaluation seconds per episode for offline methods and seconds per episode otherwise. Bold indicates the highest SR and fewest mean turns within each backbone.
Qwen3.5-4B
Qwen3.5-9B
Method
Strict-SR
Mean-Score
Turns
Time GPU·h / s
Strict-SR
Mean-Score
Turns
Time GPU·h / s
Base
17.8
25.9
39.5
87
17.2
27.8
37.5
166
Offline
MemP (offline)
33.2 ± 1.3
43.7 ± 2.0
30.0
6.9 / 74
26.6 ± 1.3
40.6 ± 2.8
31.9
12.3 / 148
ACE (offline)
5.8 ± 0.5
21.4 ± 1.1
42.0
9.4 / 122
13.2 ± 0.6
22.4 ± 1.4
41.6
16.4 / 229
EvoSkill
20.0 ± 1.6
29.3 ± 1.9
38.0
2.06 / 84
20.2 ± 2.1
33.4 ± 1.4
38.2
2.33 / 169
Table 2 : WebShop results. Strict-SR (%) denotes task_score =1 , and Mean-Score (%) averages task_score over each 500-task stream. Offline runtime is training GPU-hours / evaluation seconds per episode, and other entries are seconds per episode. Bold indicates the highest Strict-SR and Mean-Score and the fewest turns within each backbone.
Figure 3 : Cross-scene transfer and continued OaTTT on ALFWorld unseen. Transfer freezes the seen-stream adaptation state, whereas continued OaTTT resumes online updates. More in Tables 4 – 5 .
Figure 4 : Quantitative analysis on Alfworld. Columns show the base, two representative in-context adaptation failures (irrelevant or contradictory retrieved memory, or relevant guidance left unfollowed), and ASCENT, which carries LoRA updates from two earlier cool-then-place tasks and retrieves no memory. Numbers are turn indices. Checks mark valid actions, crosses mark invalid or non-action outputs, and ellipses omit turns.
Figure 5 : Analysis of online gains, preliminary co-evolution, and divergence ablation. (a) Mean success-rate gain and turns saved by ASCENT over the base, by task range and over all tasks (lighter bars 4B, darker 9B, with turns on tasks solved by both the un-evolved base model and ASCENT in Table 8 ). (b) Preliminary in-context and parametric co-evolution. (c) Distillation-divergence ablation, pre-update SR (%) at both scales, with forward KL as default.
Qwen3.5-4B
Qwen3.5-9B
Privileged information zi
Filter
SR (%)
Turns
Valid (%)
SR (%)
Turns
Valid (%)
Base
–
46.4
35.1
70.1
55.0
32.6
78.7
✗
48.8
32.9
64.4
42.9
34.1
66.2
Action only: ai,t
✓
48.3
33.5
61.2
48.6
32.3
71.4
✗
62.9
28.0
74.6
77.9
21.1
86.6
Reasoning + action + observation: (wi,t,oi,t+1)
✓
64.5
28.3
79.5
77.1
23.8
88.2
Table 3: Privileged-information ablation on ALFWorld seen, with the action-validity filter off (✗) or on (✓). The shaded row is ASCENT’s default (reasoning + action with the filter).
Figure 6 : Teacher prompt template. For a verifier-accepted episode, the privileged information zi is prepended to each shared student prefix. The empty think block reflects disabled thinking mode. Each response’s reasoning is the visible ReAct rationale before the action tag.
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg.SR
Mean turns (all)
Time s/ep
Qwen3.5-4B
Base
87.5
38.9
29.0
8.7
28.6
76.5
43.3
36.2
107
Offline
MemP (offline)
91.7
72.2
64.5
8.7
52.4
76.5
60.4 ± 4.5
30.0
69
ACE (offline)
79.2
33.3
41.9
8.7
42.9
23.5
39.6 ± 3.2
38.5
106
EvoSkill
83.3
40.7
29.0
5.8
47.6
58.8
43.0 ± 4.6
37.4
146
Table 4: ALFWorld held-out transfer. Success rate (SR, %) by task type and over the full held-out stream, mean turns per episode, and runtime (s/episode). Seen-stream artifacts remain fixed. Bold indicates the highest SR (including ties) and fewest mean turns within each backbone.
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg.SR ( Δ vs Transfer)
Mean turns (all)
Time s/ep
Qwen3.5-4B
Base
87.5
38.9
29.0
8.7
28.6
76.5
43.3
36.2
107
Online adaptation on unseen stream
MemP
100.0
94.4
74.2
21.7
33.3
64.7
64.9 ± 4.8 ( + 7.4)
28.2
58
ReasoningBank
87.5
33.3
48.4
26.1
52.4
52.9
50.7 ± 7.3 ( + 5.2)
35.0
89
ACE
75.0
11.1
3.2
0.0
9.5
35.3
21.6 ± 5.2 ( + 1.5)
45.5
225
Table 5: ALFWorld continued OaTTT. Pre-update success rate (SR, %) by task type and over the held-out stream, mean turns per episode, and runtime (s/episode). Δ is the change from frozen transfer in Table 4 . Bold indicates the highest SR (including ties) and fewest mean turns within each backbone.
Test-Normal
Test-Challenge
Method
TGC
SGC
TGC
SGC
Base
20.2
5.4
15.8
5.8
MemP
22.0 ± 2.1
1.8 ± 0.2
16.1 ± 3.3
5.8 ± 0.7
ReasoningBank
19.6 ± 1.9
3.6 ± 0.7
19.2 ± 2.8
4.3 ± 1.1
ACE
22.0 ± 4.1
3.6 ± 1.2
14.4 ± 3.9
2.2 ± 2.0
A-Mem
14.3 ± 0.4
8.9 ± 1.2
11.3 ± 1.1
2.2 ± 0.4
Table 6: AppWorld online results. Task-goal completion (TGC, %) and scenario-goal completion (SGC, %) on Test-Normal and Test-Challenge with Qwen3.5-4B.
Teacher
Schedule
SR (%)
Early → late SR (%)
Fixed
ASCENT (frozen π0 )
fixed
69.5
59.0 → 75.2
EMA of the student
EMA
β=0.999
61.4
48.6 → 77.1
EMA
β=0.99
57.1
54.3 → 57.1
EMA
β=0.90
24.3
48.6 → 2.9
Table 7: Analysis of teacher evolution in OaTTT on ALFWorld seen. The snapshot teacher is a copy of the student that is refreshed every N updates (the frozen π0 before the first refresh). Early and late refer to tasks 1–35 and 106–140, respectively.
Figure 7 : Dynamics of teacher evolution in OaTTT on ALFWorld seen. Qwen3.5-4B with evolving teachers: (a,b) trailing 30-episode success against the frozen teacher in the same setup, with triangles at snapshot refreshes, and (c) stream success against the mean teacher–student adapter distance ∥ϕˉ−ϕ∥2 before each update, with the frozen teacher as the dashed line.
Scale
Both won
ASCENT only
Base only
Mean ΔT
Median ΔT
ASCENT uses fewer turns (%)
4B
61.7± 1.2
35.7± 6.6
3.3± 1.2
−4.8± 0.9
−2.3± 0.5
63.8± 1.8
9B
74.3± 0.5
34.0± 2.2
2.7± 0.5
−7.0± 0.3
−3.5± 0.4
67.3± 4.7
Table 8: Shared-success turn comparison on ALFWorld seen. ΔT=TASCENT−TBase on tasks solved by both the un-evolved base model and ASCENT, so negative values favor ASCENT. The first three columns count the tasks in the 140-task stream solved by both, by ASCENT only, and by the base only. The last column is the percentage of shared successes on which ASCENT uses fewer turns than the base. Entries are mean ± std of per-run statistics over the 3 paired runs per scale.
Figure 8 : Memory retrieval and success on ALFWorld seen. For the retrieval-based in-context adaptation baseline, bars show means with std error bars: (a) retrieval of an available prior same-type success, and (b) success by memory availability and retrieval.
aSchool of Artificial Intelligence, Jilin University. · bEngineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University. · cInternational Center of Future Science, Jilin University. +3