A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT
Figures & tables
Figure 1 : Imitation and reinforcement lose executable behavior. Online per-episode updates (Sec. 3.2 ) of Qwen3.5-4B on the ALFWorld seen stream, with std shading: (a) first-response-token entropy, (b) KL to the initial policy, (c) valid-action rate, and (d) generated tokens per turn.
Figure 2 : Online Agentic Test-Time Training of ASCENT. The student makes a single attempt per task. A verified trajectory becomes privileged information for the frozen teacher, whose soft targets update the LoRA fast weights for the next task. Charts on the right show ALFWorld seen-stream means of 4B and 9B runs (per-scale in Fig. 5 (a) and Table 1 ). Episode time includes the update.
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg. SR
Turns
Time GPU·h / s
Qwen3.5-4B
Base
91.4
46.2
40.7
6.2
8.0
54.2
46.4
35.1
97
Offline
MemP
94.3
53.8
77.8
0.0
36.0
66.7
61.4 ± 6.5
28.6
5.2 / 61
ACE
85.7
46.2
70.4
6.2
60.0
58.3
60.7 ± 9.2
31.1
6.3 / 81
EvoSkill
93.3
17.9
35.8
6.2
32.0
66.7
49.8 ± 4.9
34.4
5.9 / 123
Table 1: ALFWorld seen-stream results. Success rate (SR, %) by task type and over the full stream, mean turns per episode, and runtime, given as training GPU-hours / evaluation seconds per episode for offline methods and seconds per episode otherwise. Bold indicates the highest SR and fewest mean turns within each backbone.
Qwen3.5-4B
Qwen3.5-9B
Method
Strict-SR
Mean-Score
Turns
Time GPU·h / s
Strict-SR
Mean-Score
Turns
Time GPU·h / s
Base
17.8
25.9
39.5
87
17.2
27.8
37.5
166
Offline
MemP (offline)
33.2 ± 1.3
43.7 ± 2.0
30.0
6.9 / 74
26.6 ± 1.3
40.6 ± 2.8
31.9
12.3 / 148
ACE (offline)
5.8 ± 0.5
21.4 ± 1.1
42.0
9.4 / 122
13.2 ± 0.6
22.4 ± 1.4
41.6
16.4 / 229
EvoSkill
20.0 ± 1.6
29.3 ± 1.9
38.0
2.06 / 84
20.2 ± 2.1
33.4 ± 1.4
38.2
2.33 / 169
Table 2 : WebShop results. Strict-SR (%) denotes task_score =1 , and Mean-Score (%) averages task_score over each 500-task stream. Offline runtime is training GPU-hours / evaluation seconds per episode, and other entries are seconds per episode. Bold indicates the highest Strict-SR and Mean-Score and the fewest turns within each backbone.
Figure 3 : Cross-scene transfer and continued OaTTT on ALFWorld unseen. Transfer freezes the seen-stream adaptation state, whereas continued OaTTT resumes online updates. More in Tables 4 – 5 .
Figure 4 : Quantitative analysis on Alfworld. Columns show the base, two representative in-context adaptation failures (irrelevant or contradictory retrieved memory, or relevant guidance left unfollowed), and ASCENT, which carries LoRA updates from two earlier cool-then-place tasks and retrieves no memory. Numbers are turn indices. Checks mark valid actions, crosses mark invalid or non-action outputs, and ellipses omit turns.
Figure 5 : Analysis of online gains, preliminary co-evolution, and divergence ablation. (a) Mean success-rate gain and turns saved by ASCENT over the base, by task range and over all tasks (lighter bars 4B, darker 9B, with turns on tasks solved by both the un-evolved base model and ASCENT in Table 8 ). (b) Preliminary in-context and parametric co-evolution. (c) Distillation-divergence ablation, pre-update SR (%) at both scales, with forward KL as default.
Qwen3.5-4B
Qwen3.5-9B
Privileged information zi
Filter
SR (%)
Turns
Valid (%)
SR (%)
Turns
Valid (%)
Base
–
46.4
35.1
70.1
55.0
32.6
78.7
✗
48.8
32.9
64.4
42.9
34.1
66.2
Action only: ai,t
✓
48.3
33.5
61.2
48.6
32.3
71.4
✗
62.9
28.0
74.6
77.9
21.1
86.6
Reasoning + action + observation: (wi,t,oi,t+1)
✓
64.5
28.3
79.5
77.1
23.8
88.2
Table 3: Privileged-information ablation on ALFWorld seen, with the action-validity filter off (✗) or on (✓). The shaded row is ASCENT’s default (reasoning + action with the filter).
Figure 6 : Teacher prompt template. For a verifier-accepted episode, the privileged information zi is prepended to each shared student prefix. The empty think block reflects disabled thinking mode. Each response’s reasoning is the visible ReAct rationale before the action tag.
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg.SR
Mean turns (all)
Time s/ep
Qwen3.5-4B
Base
87.5
38.9
29.0
8.7
28.6
76.5
43.3
36.2
107
Offline
MemP (offline)
91.7
72.2
64.5
8.7
52.4
76.5
60.4 ± 4.5
30.0
69
ACE (offline)
79.2
33.3
41.9
8.7
42.9
23.5
39.6 ± 3.2
38.5
106
EvoSkill
83.3
40.7
29.0
5.8
47.6
58.8
43.0 ± 4.6
37.4
146
Table 4: ALFWorld held-out transfer. Success rate (SR, %) by task type and over the full held-out stream, mean turns per episode, and runtime (s/episode). Seen-stream artifacts remain fixed. Bold indicates the highest SR (including ties) and fewest mean turns within each backbone.
Method
Pick
Look
Clean
Heat
Cool
Pick2
Avg.SR ( Δ vs Transfer)
Mean turns (all)
Time s/ep
Qwen3.5-4B
Base
87.5
38.9
29.0
8.7
28.6
76.5
43.3
36.2
107
Online adaptation on unseen stream
MemP
100.0
94.4
74.2
21.7
33.3
64.7
64.9 ± 4.8 ( + 7.4)
28.2
58
ReasoningBank
87.5
33.3
48.4
26.1
52.4
52.9
50.7 ± 7.3 ( + 5.2)
35.0
89
ACE
75.0
11.1
3.2
0.0
9.5
35.3
21.6 ± 5.2 ( + 1.5)
45.5
225
Table 5: ALFWorld continued OaTTT. Pre-update success rate (SR, %) by task type and over the held-out stream, mean turns per episode, and runtime (s/episode). Δ is the change from frozen transfer in Table 4 . Bold indicates the highest SR (including ties) and fewest mean turns within each backbone.
Test-Normal
Test-Challenge
Method
TGC
SGC
TGC
SGC
Base
20.2
5.4
15.8
5.8
MemP
22.0 ± 2.1
1.8 ± 0.2
16.1 ± 3.3
5.8 ± 0.7
ReasoningBank
19.6 ± 1.9
3.6 ± 0.7
19.2 ± 2.8
4.3 ± 1.1
ACE
22.0 ± 4.1
3.6 ± 1.2
14.4 ± 3.9
2.2 ± 2.0
A-Mem
14.3 ± 0.4
8.9 ± 1.2
11.3 ± 1.1
2.2 ± 0.4
Table 6: AppWorld online results. Task-goal completion (TGC, %) and scenario-goal completion (SGC, %) on Test-Normal and Test-Challenge with Qwen3.5-4B.
Teacher
Schedule
SR (%)
Early → late SR (%)
Fixed
ASCENT (frozen π0 )
fixed
69.5
59.0 → 75.2
EMA of the student
EMA
β=0.999
61.4
48.6 → 77.1
EMA
β=0.99
57.1
54.3 → 57.1
EMA
β=0.90
24.3
48.6 → 2.9
Table 7: Analysis of teacher evolution in OaTTT on ALFWorld seen. The snapshot teacher is a copy of the student that is refreshed every N updates (the frozen π0 before the first refresh). Early and late refer to tasks 1–35 and 106–140, respectively.
Figure 7 : Dynamics of teacher evolution in OaTTT on ALFWorld seen. Qwen3.5-4B with evolving teachers: (a,b) trailing 30-episode success against the frozen teacher in the same setup, with triangles at snapshot refreshes, and (c) stream success against the mean teacher–student adapter distance ∥ϕˉ−ϕ∥2 before each update, with the frozen teacher as the dashed line.
Scale
Both won
ASCENT only
Base only
Mean ΔT
Median ΔT
ASCENT uses fewer turns (%)
4B
61.7± 1.2
35.7± 6.6
3.3± 1.2
−4.8± 0.9
−2.3± 0.5
63.8± 1.8
9B
74.3± 0.5
34.0± 2.2
2.7± 0.5
−7.0± 0.3
−3.5± 0.4
67.3± 4.7
Table 8: Shared-success turn comparison on ALFWorld seen. ΔT=TASCENT−TBase on tasks solved by both the un-evolved base model and ASCENT, so negative values favor ASCENT. The first three columns count the tasks in the 140-task stream solved by both, by ASCENT only, and by the base only. The last column is the percentage of shared successes on which ASCENT uses fewer turns than the base. Entries are mean ± std of per-run statistics over the 3 paired runs per scale.
Figure 8 : Memory retrieval and success on ALFWorld seen. For the retrieval-based in-context adaptation baseline, bars show means with std error bars: (a) retrieval of an available prior same-type success, and (b) success by memory availability and retrieval.
LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies that previously worked. Test-time training (TTT) offers a way to adapt model weights to the evolving task state, but existing LLM TTT methods largely adapt once to a fixed input. We study continuous TTT in multi-turn agent episodes, where each update changes the policy that generates later training text. This creates a self-training loop that helps when new trajectory information appears, but can amplify drift when the agent gets stuck and repeatedly trains on similar text. We find that update-text repetition distinguishes these regimes and introduce Agentic Test-Time Training (aTTT), a token-level reweighting method that downweights the loss on tokens appearing in repeated n-grams from prior updates while leaving novel tokens fully weighted. To run such updates inside live episodes, we build a concurrent serving system using vLLM's runtime LoRA API, limiting overhead to 1.9× the no-TTT cost. aTTT improves success by up to 5.0 points on ALFWorld and 4.9 points on SWE-bench Lite. The gains concentrate where models already have task competence but drift over long trajectories, suggesting that aTTT mainly preserves existing competence rather than teaching new abilities.
Test-time training (TTT) adapts an LLM's parameters using signals derived from test inputs, and can make striking improvements in pre-specified settings such as IMO competitions or designated open problems. By turning deployment experience into parameter updates, TTT provides a direct mechanism for model-level self-improvement. Yet TTT is not universally beneficial: each TTT algorithm works in different settings, and applying an ill-suited method could waste test-time compute or even damage model performance. Therefore, such parameter-level self-improvement requires agency: the model must decide when TTT is warranted, which algorithm to invoke, and whether an existing skill can be reused. To fill this gap, we introduce Agentic-TTT, which learns a test-time policy to govern those decisions. Agentic-TTT turns TTT procedures into callable tools, treats accumulated skills as an evolving deployment environment, and trains its policy using the observed utility gains from its decisions. On our benchmark, Agentic-TTT nearly doubles the utility over the backbone model, learns to trade off utility against compute, and generalizes to domains unseen during training. Together, these results point toward autonomous self-improvement: models that can decide how to learn from their own deployment experience.
Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparametric framework that separates immediate use from persistent trust. It turns informative transitions into hypotheses that can guide the next step, while requiring prospective validation before reuse across episodes. Their predicted effects are checked against subsequent observations outside the source episodes, and only sufficiently supported hypotheses become available for persistent guidance. This process updates external knowledge while keeping all model parameters fixed. Over five rounds on WebArena-Lite and ALFWorld, StepLearn achieves average success rates of 59.9% and 84.0% with GPT-5-mini, and 57.8% and 88.1% with Qwen3.5-35B-A3B, respectively. It outperforms EvoTest, the strongest baseline, by 2.2-12.7 percentage points across the four settings. Learning dynamics further shows that these gains are not restricted to the final repetition, with advantages already present on first task attempts in most settings.