While Large Language Model (LLM) agents excel at general tasks, they inherently struggle with continual adaptation due to the frozen weights after deployment. Conventional reinforcement learning (RL) offers a solution but incurs prohibitive computational costs and the risk of catastrophic forgetting. We introduce Just-In-Time Reinforcement Learning (JitRL), a training-free framework that enables test-time policy optimization without any gradient updates. JitRL maintains a dynamic, non-parametric memory of experiences and retrieves relevant trajectories to estimate action advantages on-the-fly. These estimates are then used to directly modulate the LLM's output logits. We theoretically prove that this additive update rule is the exact closed-form solution to the KL-constrained policy optimization objective. Extensive experiments on WebArena and Jericho demonstrate that JitRL establishes a new state-of-the-art among training-free methods. Crucially, JitRL outperforms the performance of computationally expensive fine-tuning methods (e.g., WebRL) while reducing monetary costs by over 30 times, offering a scalable path for continual learning agents. The code is available at https://github.com/liushiliushi/JitRL.
Figures & tables
Figure 1 : While standard RL performs policy gradient updates during training using previous trajectories, JitRL operates at test time. Specifically, it retrieves trajectories relevant to the current state to estimate advantages A , subsequently refining the output logits through a KL-regularized policy optimization objective.
Figure 2 : Overview of the Just-In-Time Reinforcement Learning (JitRL) framework. The system operates in a continuous loop: (1) In the Inference (top), the agent retrieves relevant past experiences N(s) from the non-parametric memory M . The base LLM’s logits z are then adjusted in closed-form ( z′=z+βA ) using the estimated advantage A derived from historical returns, enabling test-time policy improvement without gradient updates. (2) In the Memory Update (bottom), completed trajectories are analyzed by an evaluator to compute discounted returns Gt . These new experiences are stored back into M , allowing the agent to evolve its policy across episodes.
Method
Library
Zork1
Zork3
Avg
Final
Avg
Final
Avg
Final
Static
10.0
10
8.5
10
0.2
0
Memory
13.2
14
22.9
25
1.0
1
Reflexion
15.4
18
26.1
35
1.4
1
AWM
12.7
10
38.8
44
1.9
2
EvoTest
21.5
26
46.8
54
2.6
4
Table 3 : Results on Jericho games. We report: Avg : average score across 50 episodes; Final : final episode score.
Method
Admin
Reddit
Avg
Final
Avg
Final
Gemini-2.5-flash
Static
39.46
39.67
39.60
42.86
Memory
47.91
47.80
53.02
55.04
Reflexion
48.46
50.55
54.88
56.59
AWM
49.46
51.09
55.66
58.14
Table 4 : Comparison of average ( Avg ) and Final success rate between JitRL and baseline methods across different backbone LLMs.
Task
Candidate Action
Base
JitRL
Mechanism Explanation
Find customer reviews
click(CATALOG)
0.90
0.40
While “Catalog” appears semantically intuitive, memory corrects this prior: reviews are located under “Marketing”.
(Site Functionality)
click(MARKETING)
0.70
1.40
Find posts in subreddit
fill(Search, “..”)
0.95
0.45
Global search often yields noisy results. Memory steers the agent toward “Forums” for a deterministic and accurate path.
(Navigation Precision)
click(Forums)
0.80
1.80
Access product inventory
click(Products)
0.95
-0.28
Clicking loads a generic page. Memory identifies “hover” as an efficient shortcut to instantly reveal subcategories.
(UI Mechanics)
hover(Products)
0.40
0.90
Table 7 : Qualitative Analysis of Policy Improvement. We select three representative cases showing how JitRL modifies decision-making. Base and JitRL denote the logits before and after the memory-based update. Compared to Base , JitRL decreases the logits for incorrect options and increases them for correct options. In this way, JitRL corrects semantic priors and optimizes efficiency.
Method
Admin
Reddit
JitRL (Prompt Update)
48.35
54.57
JitRL (Logit Update)
52.31
57.64
Table 8 : Ablation: Logit update vs. prompting on WebArena.
Metric
Static
Memory
Reflexion
AWM
EvoTest
WebRL
JitRL
Cost
$200
$230
$220
$250
$220
∼ $9900
$290
Table 9 : Comparison of training cost.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Parameter
Value
Environment Settings
Step limit per episode
60
Number of episodes
50
LLM Configuration
Temperature
0.8
Candidate actions ( ∣C∣ )
3
Appendix
Table 10 : Hyperparameters for Jericho experiments.
Category
Parameter
Value
Environment Settings
Step limit per episode
10
Number of episodes
50
Random seed
0
LLM Configuration
Temperature
0.8
Appendix
Table 11 : Hyperparameters for WebArena experiments.
Method
Admin
GitLab
Map
Reddit
Shopping
Avg
WebRL
38.89
23.53
9.68
50.00
17.39
27.27
JitRL
28.57
30.00
20.00
52.63
39.13
32.97
Appendix
Table 12 : Final success rate comparison (%) between WebRL and JitRL on WebArena-Lite using the same base model and training data.
Method
Admin
GitLab
Map
Reddit
Shopping
Average
Cost
Llama-3.1-70B
10.50
16.70
17.10
20.00
4.40
12.70
–
SFT
20.00
20.00
26.70
52.60
13.30
23.00
$640
JitRL
47.22
38.24
25.81
58.33
34.78
40.88
$200
WebRL
58.33
47.06
32.26
62.50
30.43
46.06
$9,900
Appendix
Table 13 : Controlled comparison on WebArena-Lite (Final success rate %) using the same Llama-3.1-70B-Instruct backbone, with separate training and evaluation phases.
Method
Library
Zork1
Zork3
GRPO
13.6
16.2
1.1
JitRL
20.8
42.1
2.3
Appendix
Table 14 : Same-backbone comparison on Jericho (Qwen3-32B, both methods adapt on the test set, Avg score over 50 episodes).
Method
Library
Zork1
Zork3
EvoTest (best baseline)
21.5
46.8
2.6
JitRL (unified)
24.7
51.8
3.3
JitRL
25.9
53.0
3.1
Appendix
Table 15 : Unified pipeline comparison on Jericho (Avg score over 50 episodes).
Memory encodes that this action immediately yields +5 points and unlocks the rare books room.
Zork1: Loud Room
take platinum bar
0.92
0.42
Greedy treasure-grabbing seems intuitive, but fails
(Deafening noise)
echo
0.30
1.50
in the noisy room. JitRL learns that “echo” quiets the room, enabling treasure collection.
Zork3: Cliff
climb down
0.90
0.30
Direct descent without preparation leads to death.
(Holding rope)
tie rope to railing
0.40
1.60
Memory encodes that securing the rope first enables safe descent and progression.
Appendix
Table 17 : Qualitative Analysis of Policy Improvement on Jericho. We select one representative case from each text-based game showing how JitRL modifies decision-making. Base and JitRL denote the logits before and after the memory-based update. The results highlight JitRL’s ability to override exploratory defaults, learn game-specific mechanics, and prioritize high-reward actions.
Website
GitLab
Map
Reddit
Shopping
Admin
Overall
Avg Steps
5.83
9.39
4.57
5.55
6.20
6.28
Appendix
Table 18 : Average Steps per Website.
Memory Size (# entries)
Retrieval (ms)
Avg Score
0–500
15–22
18.1
500–1000
26–27
25.1
1000–1500
29
27.2
1500–2000
38
29.2
2000–2500
47
30.0
Appendix
Table 19 : Inference overhead and performance as memory grows (Library, 50 episodes).
Rollout
Batch Size
PPO Epochs
LR
Val Mean
Val Max
Val Min
8
1
1
1e-6
16.2
35
0
16
1
1
1e-6
13.6
40
0
32
1
1
1e-6
12.0
40
-5
8
1
2
1e-6
10.4
40
0
8
1
4
1e-6
17.1
40
5
8
1
4
1e-5
25.5
45
-5
Appendix
Table 20 : GRPO hyperparameter sweep on Zork1 (Qwen3-32B, 50 training steps).
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China · The Beijing University of Posts and Telecommunications, Beijing, China