LLM agents increasingly improve at inference time by accumulating experience in context rather than by updating parameters. This process is often described as in-context reinforcement learning (ICRL). Whether in-context learning (ICL) can actually play the role of RL, however, has not been tested. We study this question in its simplest form, direct ICRL, where the model conditions directly on raw trajectory-reward pairs, and ask whether the reward acts as a learning signal. Through controlled experiments on four benchmarks across six models, we find that the reward is read, but its effect is small: flipping, randomizing, or removing the reward leaves the improvement curve almost unchanged, and this holds even under meta-prompts that explicitly instruct the model to explore, exploit, or reason over rewards. Trajectories drive improvement, but not through their semantic content: shuffled or corrupted trajectories work as well as real ones. These patterns closely mirror those known in ICL, suggesting that direct ICRL is better understood as a special case of ICL than as inference-time RL. This reframing has implications for agent memory design: ICL factors such as input distribution and demonstrations may matter more than RL elements such as reward shaping and exploration.
Figures & tables
Figure 1 : Direct ICRL overview (left) and controlled interventions for probing its mechanism: reward (middle) and trajectory perturbations (right).
Figure 2 : Episode-level performance under different reward conditions on three benchmarks.
Model
No Imp.
Reward-Agnostic
Reward-Sensitive
Total
Qwen-3-4B-Instruct
25 (83.3%)
4 (13.3%)
1 (3.3%)
30
Llama-3.1-8B-Instruct
27 (90.0%)
2 (6.7%)
1 (3.3%)
30
Qwen-3.5-9B
19 (63.3%)
10 (33.3%)
1 (3.3%)
30
GPT-5-nano
11 (36.7%)
19 (63.3%)
0 (0.0%)
30
GPT-5-mini
7 (23.3%)
22 (73.3%)
1 (3.3%)
30
Gemini-3.1-Flash-Lite
12 (40.0%)
16 (53.3%)
2 (6.7%)
30
Table 1: Comparison of improvement types across models on HumanEval+ and MBPP+. Each entry reports the number of problems and the percentage within each model-benchmark setting.
ΔLOO
Target Replay Rate
Memory Replay Rate
Condition
τ−
τ+
τ−
τ+
Problem-only
–
–
0.7
1.1
16.3
Trajectory-only
0.416
0.162
3.1
1.8
34.3
Table 2 : Effect of showing trajectories, without any reward. Replay rates are in %. ΔLOO is undefined for Problem-only because no memory is shown.
ΔLOO
Target Replay Rate
Condition
τ−
τ+
τ−
τ+
Trajectory-only (no reward)
0.416
0.162
3.1
1.8
ICRL (true reward)
0.360
0.173
2.5
3.0
Specific-flip
0.474
0.145
3.0
1.6
Table 3 : Effect of the reward field on the replay of the target trajectory. Trajectories are identical across rows; only the reward field differs. Target Replay Rate is in %.
Target Replay Rate
Memory Replay Rate
Change
Condition
τ−
τ+
τ−
τ+
Reference
Problem-only
0.7
1.1
16.3
16.3
ICRL
2.5
3.0
30.5
30.5
Position
Order-first
2.4
5.5
29.3
30.0
Order-last
3.2
13.2
29.4
41.9
Discouraging replay
All Failure
2.3
1.7
27.6
27.3
Table 4 : Effect of changes other than the target’s reward. All replay rates are in %. The τ− and τ+ columns report runs in which the target is the lowest- and the highest-reward trajectory, respectively. Problem-only is the level without any memory.
Memory Setting
MBPP
ScienceWorld
RSM
0.250
-32.479
Cross-Task
0.122
-43.000
Random Valid Action
0.102
-68.208
ICRL History
0.245
-42.238
Trajectory Perturbation
Shuffle
0.243
-38.588
Table 5 : Mean reward across MBPP and ScienceWorld. Standard deviations are reported in Table 9 .
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Benchmark
Comparison
Mean Δ
95% CI
p
ScienceWorld
True vs. Random
−0.76
[−9.40,7.86]
0.86
True vs. No Reward
+1.96
[−6.73,10.64]
0.65
MBPP+
True vs. Random
+0.028
[−0.042,0.098]
0.43
True vs. No Reward
+0.014
[−0.064,0.091]
0.72
HumanEval+
True vs. Random
+0.005
[−0.064,0.075]
0.88
True vs. No Reward
+0.037
[−0.056,0.130]
0.42
Appendix
Table 6 : Task-level paired t -tests on final episode scores ( n=40 tasks per benchmark, paired across 3 seeds). All six confidence intervals include zero, and no comparison reaches p<0.05 . ScienceWorld final scores span roughly 80 points; coding scores are unit-test accuracy in [0,1] .
Figure 3 : Improvement category for accumulated memory settings.
Figure 4 : ScienceWorld ICRL experiment using Llama-3.1-8B-Instruct as a policy network.
Figure 5 : ScienceWorld ICRL experiment with verbal reward.
Figure 6 : ScienceWorld ICRL experiment with Reflexion.
Figure 7 : Top-k and Bottom-k memory selection on MBPP+ and HumanEval+.
ΔLOO
Method Type
τ−
τ+
ICRL
True Reward
0.3601
0.1734
All Failure
0.4066
0.1468
Anti prompt
0.3322
0.1632
CoT
0.4339
0.1615
Duplicate
0.4327
0.2504
Appendix
Table 7 : Full results of ΔLOO
Type
Detail
Target Replay Rate (%)
Memory Replay Rate (%)
τ+
τ−
τ+
τ−
Problem Only
Problem Only
1.1
0.7
16.3
16.3
ICRL
True Reward
3.0
2.5
30.5
30.5
All Failure
1.7
2.3
27.3
27.6
Specific Failure
1.7
2.2
30.6
30.5
Anti Prompt
2.3
2.4
30.0
30.2
Appendix
Table 8 : Full replay-rate results across memory conditions.
Memory Setting
MBPP
ScienceWorld
RSM
0.250±0.021
−32.479±4.520
Cross-Task
0.122±0.017
−43.000±4.100
Random Valid Action
0.102±0.012
−68.208±3.880
ICRL History
0.245±0.032
−42.238±2.860
Trajectory Perturbation
Shuffle
0.243±0.023
−38.588±6.090
Appendix
Table 9 : Mean reward comparison across MBPP and ScienceWorld with standard deviations.
Figure 8 : Example of the Crossover perturbation on an MBPP+ task. Two parent trajectories (a, b), each a valid solution sampled by the model, are mixed line-by-line into a structurally broken child (c). Even though the child code does not compile, using such perturbed trajectories as in-context memory yields performance comparable to using real ICRL trajectories ( Section 6 , Table 9 ).
Figure 9 : Example of the Shuffle perturbation on a ScienceWorld task. The first and last actions are kept fixed; the middle three steps are randomly reordered. The shuffled version is no longer a sensible plan — focus on potato is repeated immediately at step 2, and the agent moves the potato to the red box at step 5 without re-focusing on it — yet using such trajectories as in-context memory yields performance comparable to using the original trajectories ( Section 6 ).