Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparametric framework that separates immediate use from persistent trust. It turns informative transitions into hypotheses that can guide the next step, while requiring prospective validation before reuse across episodes. Their predicted effects are checked against subsequent observations outside the source episodes, and only sufficiently supported hypotheses become available for persistent guidance. This process updates external knowledge while keeping all model parameters fixed. Over five rounds on WebArena-Lite and ALFWorld, StepLearn achieves average success rates of 59.9% and 84.0% with GPT-5-mini, and 57.8% and 88.1% with Qwen3.5-35B-A3B, respectively. It outperforms EvoTest, the strongest baseline, by 2.2-12.7 percentage points across the four settings. Learning dynamics further shows that these gains are not restricted to the final repetition, with advantages already present on first task attempts in most settings.
Figures & tables
Figure 1 : Learning during interaction. Unlike episode-level learning, StepLearn acquires knowledge for the next decision and validates it in later episodes before persistent reuse.
Figure 2 : StepLearn overview. (a) Knowledge acquisition turns informative transitions into policy–effect hypotheses. (b) Immediate reuse retrieves runtime and verified knowledge for action selection. (c) Prequential verification fixes predictions before execution and accumulates evidence from non-source episodes for persistent reuse. The example is schematic: verdicts are categories, not counts, and both verified-memory icons denote one store.
Method
WebArena-Lite
ALFWorld
Admin
GitLab
Map
Reddit
Shop.
Avg.
Look
Pick
Clean
Cool
Heat
Two
Avg.
GPT-5-mini
Static
51.4
50.0
23.1
78.9
40.0
46.5
25.0
20.8
19.4
28.6
17.4
20.6
21.6
Memory
57.4
43.3
19.2
75.8
42.7
45.8
75.6
20.8
38.7
68.6
27.0
21.2
40.8
Reflexion
59.3
43.3
23.1
80.0
44.8
47.7
52.2
44.2
52.9
46.7
57.4
58.8
51.8
JitRL
41.7
57.3
24.6
65.3
38.7
43.9
66.7
78.3
45.8
36.2
43.5
51.8
53.3
Table 1 : Average success rates (%). Bold and underline mark the best and second best results per column within each model block.
Figure 3 : Success rates over five task repetitions (rounds), with panel-specific vertical scales. Qwen3.5-35B denotes Qwen3.5-35B-A3B in thinking mode.
Figure 4 : First-attempt (R1) success rates. R1 may include knowledge from earlier tasks. Bars follow legend order.
Configuration
GPT-5-mini
Qwen3.5-35B-A3B
WebArena-Lite
ALFWorld
WebArena-Lite
ALFWorld
StepLearn (full reference)
59.9
84.0
57.8
88.1
w/o stepwise acquisition
56.6 ( −3.3 )
81.6 ( −2.4 )
49.8 ( −8.0 )
86.9 ( −1.2 )
w/o runtime reuse
56.8 ( −3.1 )
81.8 ( −2.2 )
51.2 ( −6.6 )
87.2 ( −0.9 )
w/o retention safeguards
57.0 ( −2.9 )
76.4 ( −7.6 )
52.0 ( −5.8 )
84.2 ( −3.9 )
Table 2 : Mechanism ablations: average success rates (%). Parentheses show percentage-point changes from the full reference.
Auxiliary learner
WebArena-Lite
ALFWorld
Actor: GPT-5-mini
GPT-5.4-nano (reference)
59.9
84.0
GPT-5-mini
55.9
86.4
Actor: Qwen3.5-35B-A3B
Qwen3.5-4B (reference)
57.8
88.1
Qwen3.5-9B
56.8
89.0
Table 3 : Auxiliary learners: success rates (%).
Method
WebArena-Lite
ALFWorld
Static
23.8
43.6
Memory
32.8
188.8
Reflexion
25.5
78.4
JitRL
71.4
196.4
AWM
25.4
58.8
EvoTest
85.2
61.2
Table 4 : API and training costs (USD).
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
API configuration
Local configuration
Actor
GPT-5-mini
Qwen3.5-35B-A3B
Auxiliary learner
GPT-5.4-nano
Qwen3.5-4B
Temperature
0.0
0.0
Action limit: WebArena / ALFWorld
25 / 50
25 / 50
Knowledge retrieval limit
6
6
Stagnation threshold
2 steps
2 steps
Appendix
Table 5 : Inference and knowledge settings for StepLearn .
Method
Benchmark
H20 GPU-hours
Cost (USD)
SFT
WebArena-Lite
29.8
59.6
SFT
ALFWorld
12.2
24.4
TTRL
WebArena-Lite
38.0
76.0
TTRL
ALFWorld
54.0
108.0
Appendix
Table 6 : Training resources and costs for the retained pipelines.
aSchool of Artificial Intelligence, Jilin University. · bEngineering Research Center of Knowledge-Driven Human-Machine Intelligence, Jilin University. · cInternational Center of Future Science, Jilin University. +3