While Large Language Model (LLM) agents excel at general tasks, they inherently struggle with continual adaptation due to the frozen weights after deployment. Conventional reinforcement learning (RL) offers a solution but incurs prohibitive computational costs and the risk of catastrophic forgetting. We introduce Just-In-Time Reinforcement Learning (JitRL), a training-free framework that enables test-time policy optimization without any gradient updates. JitRL maintains a dynamic, non-parametric memory of experiences and retrieves relevant trajectories to estimate action advantages on-the-fly. These estimates are then used to directly modulate the LLM's output logits. We theoretically prove that this additive update rule is the exact closed-form solution to the KL-constrained policy optimization objective. Extensive experiments on WebArena and Jericho demonstrate that JitRL establishes a new state-of-the-art among training-free methods. Crucially, JitRL outperforms the performance of computationally expensive fine-tuning methods (e.g., WebRL) while reducing monetary costs by over 30 times, offering a scalable path for continual learning agents. The code is available at https://github.com/liushiliushi/JitRL.
Figures & tables
Figure 1 : While standard RL performs policy gradient updates during training using previous trajectories, JitRL operates at test time. Specifically, it retrieves trajectories relevant to the current state to estimate advantages A , subsequently refining the output logits through a KL-regularized policy optimization objective.
Figure 2 : Overview of the Just-In-Time Reinforcement Learning (JitRL) framework. The system operates in a continuous loop: (1) In the Inference (top), the agent retrieves relevant past experiences N(s) from the non-parametric memory M . The base LLM’s logits z are then adjusted in closed-form ( z′=z+βA ) using the estimated advantage A derived from historical returns, enabling test-time policy improvement without gradient updates. (2) In the Memory Update (bottom), completed trajectories are analyzed by an evaluator to compute discounted returns Gt . These new experiences are stored back into M , allowing the agent to evolve its policy across episodes.
Method
Library
Zork1
Zork3
Avg
Final
Avg
Final
Avg
Final
Static
10.0
10
8.5
10
0.2
0
Memory
13.2
14
22.9
25
1.0
1
Reflexion
15.4
18
26.1
35
1.4
1
AWM
12.7
10
38.8
44
1.9
2
EvoTest
21.5
26
46.8
54
2.6
4
Table 3 : Results on Jericho games. We report: Avg : average score across 50 episodes; Final : final episode score.
Method
Admin
Reddit
Avg
Final
Avg
Final
Gemini-2.5-flash
Static
39.46
39.67
39.60
42.86
Memory
47.91
47.80
53.02
55.04
Reflexion
48.46
50.55
54.88
56.59
AWM
49.46
51.09
55.66
58.14
Table 4 : Comparison of average ( Avg ) and Final success rate between JitRL and baseline methods across different backbone LLMs.
Task
Candidate Action
Base
JitRL
Mechanism Explanation
Find customer reviews
click(CATALOG)
0.90
0.40
While “Catalog” appears semantically intuitive, memory corrects this prior: reviews are located under “Marketing”.
(Site Functionality)
click(MARKETING)
0.70
1.40
Find posts in subreddit
fill(Search, “..”)
0.95
0.45
Global search often yields noisy results. Memory steers the agent toward “Forums” for a deterministic and accurate path.
(Navigation Precision)
click(Forums)
0.80
1.80
Access product inventory
click(Products)
0.95
-0.28
Clicking loads a generic page. Memory identifies “hover” as an efficient shortcut to instantly reveal subcategories.
(UI Mechanics)
hover(Products)
0.40
0.90
Table 7 : Qualitative Analysis of Policy Improvement. We select three representative cases showing how JitRL modifies decision-making. Base and JitRL denote the logits before and after the memory-based update. Compared to Base , JitRL decreases the logits for incorrect options and increases them for correct options. In this way, JitRL corrects semantic priors and optimizes efficiency.
Method
Admin
Reddit
JitRL (Prompt Update)
48.35
54.57
JitRL (Logit Update)
52.31
57.64
Table 8 : Ablation: Logit update vs. prompting on WebArena.
Metric
Static
Memory
Reflexion
AWM
EvoTest
WebRL
JitRL
Cost
$200
$230
$220
$250
$220
∼ $9900
$290
Table 9 : Comparison of training cost.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Parameter
Value
Environment Settings
Step limit per episode
60
Number of episodes
50
LLM Configuration
Temperature
0.8
Candidate actions ( ∣C∣ )
3
Appendix
Table 10 : Hyperparameters for Jericho experiments.
Category
Parameter
Value
Environment Settings
Step limit per episode
10
Number of episodes
50
Random seed
0
LLM Configuration
Temperature
0.8
Appendix
Table 11 : Hyperparameters for WebArena experiments.
Method
Admin
GitLab
Map
Reddit
Shopping
Avg
WebRL
38.89
23.53
9.68
50.00
17.39
27.27
JitRL
28.57
30.00
20.00
52.63
39.13
32.97
Appendix
Table 12 : Final success rate comparison (%) between WebRL and JitRL on WebArena-Lite using the same base model and training data.
Method
Admin
GitLab
Map
Reddit
Shopping
Average
Cost
Llama-3.1-70B
10.50
16.70
17.10
20.00
4.40
12.70
–
SFT
20.00
20.00
26.70
52.60
13.30
23.00
$640
JitRL
47.22
38.24
25.81
58.33
34.78
40.88
$200
WebRL
58.33
47.06
32.26
62.50
30.43
46.06
$9,900
Appendix
Table 13 : Controlled comparison on WebArena-Lite (Final success rate %) using the same Llama-3.1-70B-Instruct backbone, with separate training and evaluation phases.
Method
Library
Zork1
Zork3
GRPO
13.6
16.2
1.1
JitRL
20.8
42.1
2.3
Appendix
Table 14 : Same-backbone comparison on Jericho (Qwen3-32B, both methods adapt on the test set, Avg score over 50 episodes).
Method
Library
Zork1
Zork3
EvoTest (best baseline)
21.5
46.8
2.6
JitRL (unified)
24.7
51.8
3.3
JitRL
25.9
53.0
3.1
Appendix
Table 15 : Unified pipeline comparison on Jericho (Avg score over 50 episodes).
Memory encodes that this action immediately yields +5 points and unlocks the rare books room.
Zork1: Loud Room
take platinum bar
0.92
0.42
Greedy treasure-grabbing seems intuitive, but fails
(Deafening noise)
echo
0.30
1.50
in the noisy room. JitRL learns that “echo” quiets the room, enabling treasure collection.
Zork3: Cliff
climb down
0.90
0.30
Direct descent without preparation leads to death.
(Holding rope)
tie rope to railing
0.40
1.60
Memory encodes that securing the rope first enables safe descent and progression.
Appendix
Table 17 : Qualitative Analysis of Policy Improvement on Jericho. We select one representative case from each text-based game showing how JitRL modifies decision-making. Base and JitRL denote the logits before and after the memory-based update. The results highlight JitRL’s ability to override exploratory defaults, learn game-specific mechanics, and prioritize high-reward actions.
Website
GitLab
Map
Reddit
Shopping
Admin
Overall
Avg Steps
5.83
9.39
4.57
5.55
6.20
6.28
Appendix
Table 18 : Average Steps per Website.
Memory Size (# entries)
Retrieval (ms)
Avg Score
0–500
15–22
18.1
500–1000
26–27
25.1
1000–1500
29
27.2
1500–2000
38
29.2
2000–2500
47
30.0
Appendix
Table 19 : Inference overhead and performance as memory grows (Library, 50 episodes).
Rollout
Batch Size
PPO Epochs
LR
Val Mean
Val Max
Val Min
8
1
1
1e-6
16.2
35
0
16
1
1
1e-6
13.6
40
0
32
1
1
1e-6
12.0
40
-5
8
1
2
1e-6
10.4
40
0
8
1
4
1e-6
17.1
40
5
8
1
4
1e-5
25.5
45
-5
Appendix
Table 20 : GRPO hyperparameter sweep on Zork1 (Qwen3-32B, 50 training steps).
LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies that previously worked. Test-time training (TTT) offers a way to adapt model weights to the evolving task state, but existing LLM TTT methods largely adapt once to a fixed input. We study continuous TTT in multi-turn agent episodes, where each update changes the policy that generates later training text. This creates a self-training loop that helps when new trajectory information appears, but can amplify drift when the agent gets stuck and repeatedly trains on similar text. We find that update-text repetition distinguishes these regimes and introduce Agentic Test-Time Training (aTTT), a token-level reweighting method that downweights the loss on tokens appearing in repeated n-grams from prior updates while leaving novel tokens fully weighted. To run such updates inside live episodes, we build a concurrent serving system using vLLM's runtime LoRA API, limiting overhead to 1.9× the no-TTT cost. aTTT improves success by up to 5.0 points on ALFWorld and 4.9 points on SWE-bench Lite. The gains concentrate where models already have task competence but drift over long trajectories, suggesting that aTTT mainly preserves existing competence rather than teaching new abilities.
Large language models (LLMs) often suffer from catastrophic forgetting in continual learning: after learning new tasks sequentially, they perform worse on earlier tasks. Existing methods mitigate catastrophic forgetting by data replay, parameter freezing, or regularization. However, these methods lack understanding of LLM mechanisms and cannot distinguish which parameters store important knowledge from previous tasks and which parameters can be updated for new tasks. To address this, we propose the attribution-guided continual fine-tuning framework that leverages Layer-wise Relevance Propagation (LRP) to estimate parameter importance based on the internal computational process of LLMs. During continual learning, parameters critical to previous tasks are constrained to receive smaller updates, while less relevant parameters remain available for learning new tasks. Extensive experiments show that, compared with baseline methods, our approach reduces catastrophic forgetting while preserving adaptability to new tasks, highlighting the value of mechanistic attribution for continual fine-tuning of LLMs.
Yazheng Liu, Yuxuan Wan, Rui Xu +3
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China · The Beijing University of Posts and Telecommunications, Beijing, China
We present AgentJet, a distributed swarm training framework for large language model (LLM) agent reinforcement learning. Unlike centralized frameworks that tightly couple agent rollouts with model optimization, AgentJet adopts a decoupled multi-node architecture in which swarm server nodes host trainable models and run optimization on GPU clusters, whereas swarm client nodes execute arbitrary agents on arbitrary devices. This design provides capabilities that are difficult to support in centralized frameworks: (1) heterogeneous multi-model reinforcement learning, enabling the training of heterogeneous multi-agent teams with multiple LLM as brains; (2) multi-task cocktail training with isolated agent runtimes; (3) fault-tolerant execution that prevents external environment failures from interrupting the training process; and (4) live code iteration, which allows agents to be edited during training by replacing swarm client nodes. To support efficient RL in multi-model, multi-turn, and multi-agent settings, AgentJet introduces a context tracking module with timeline merging, which consolidates redundant context and achieves a 1.5-10x training speedup. Finally, AgentJet introduces an automated research system that takes a research topic as input and autonomously conducts long-horizon, multi-day RL studies on large-scale clusters. By leveraging the swarm architecture, this system reproduces key exploratory workflows of RL researchers without human intervention during execution.