Imitation learning (IL) teaches language-model agents to reproduce expert actions but not to distinguish them from plausible mistakes. Self-reflection methods expose models to alternatives yet use supervised fine-tuning (SFT) to imitate fixed rationales and actions. We introduce Agentic Critical Training (ACT), which uses reinforcement learning with verifiable rewards (RLVR) to train models to judge actions directly. At each expert-trajectory state, ACT pairs an expert action with an alternative sampled from the initial policy and randomizes their order. The model generates its own reasoning but is rewarded only for selecting the expert action. ACT reuses demonstrations, requires no reference rationales, and allows pair reuse across model sizes. ACT is a warm-up before IL, optionally followed by RL; inference requires no candidate comparison. Across Qwen3-8B and Olmo-3-7B-Instruct on ALFWorld-ID, WebShop, and ScienceWorld, ACT yields average gains of 5.85 points over IL and 4.12 points over IL→RL without ACT, while also improving ALFWorld-OOD. With Olmo on ScienceWorld, the full pipeline gains 15.36 points over CoT prompting and 9.23 points over IL→RL without ACT. Both ACT→IL and the full pipeline outperform supervised reflection baselines. Controls with fixed pairs or matched training durations show that the gains stem from the ACT objective rather than additional data or training. Without reasoning-specific post-training, the standalone ACT checkpoint achieves the highest mean among evaluated models on MATH-500 and GPQA-Diamond, showing that action comparison complements generation.
Figures & tables
Figure 1 : ACT trains action judgment with a verifiable selection reward. (a) Early Experience rolls out both candidates and uses SFT to imitate a synthesized reflection and expert action. (b) ACT randomizes the pair and rewards expert-action selection without a reference rationale. ACT precedes IL; RL continues from IL along the upper path, and the bypass goes directly to test. Blue and green bars show mean gains from adding ACT on ALFWorld-ID, WebShop, and ScienceWorld. Purple bars compare the ALFWorld-only Qwen3-8B ACT-stage checkpoint with the original Qwen3-8B model on general reasoning. All bars share a 0–7 point scale.
Figure 2 : Overview of the ACT training pipeline. Stage 1 constructs expert–alternative pairs from existing demonstrations. Stage 2 trains action judgment with GRPO using the action-selection reward in eq. 2 . Its accuracy component verifies expert-action selection; admissibility and format terms are detailed in Section A.1 ( Radm is disabled for WebShop). Stage 3 applies IL from the ACT checkpoint, followed by RL from the resulting IL checkpoint or the bypass to test. Both paths produce a policy that generates actions directly without candidate pairs at test time.
Method
ALFWorld
WebShop
ScienceWorld
ID
OOD
Prompt w/o CoT thinking
32.62 ± 2.63
24.88 ± 2.88
18.20 ± 0.85
16.99 ± 2.85
Prompt w/ CoT thinking
50.00 ± 4.77
39.30 ± 7.72
15.53 ± 0.34
18.68 ± 0.34
ACT stage
62.14 ± 7.72
59.45 ± 9.23
21.53 ± 4.58
30.04 ± 2.13
Imitation Learning (IL)
85.71 ± 0.58
79.85 ± 2.79
28.00 ± 1.31
21.94 ± 2.65
Early Experience
86.43 ± 1.55
79.10 ± 10.57
30.93 ± 1.23
22.82 ± 1.30
Table 1 : Main results on Qwen3-8B. ALFWorld and WebShop report success rates (%); ScienceWorld reports mean episode score (0–100).
Method
ALFWorld
WebShop
ScienceWorld
ID
OOD
Prompt w/o CoT thinking
0.71 ± 0.58
0.25 ± 0.35
11.47 ± 0.34
3.47 ± 2.04
Prompt w/ CoT thinking
11.91 ± 0.67
10.94 ± 2.14
12.47 ± 0.47
6.31 ± 3.97
ACT stage
16.90 ± 1.69
11.94 ± 3.71
11.53 ± 1.31
12.03 ± 0.92
Imitation Learning (IL)
62.62 ± 3.37
56.97 ± 9.33
12.87 ± 9.02
8.10 ± 0.19
Early Experience
42.62 ± 3.51
35.32 ± 6.40
13.27 ± 7.04
8.97 ± 0.71
Table 2 : Main results on Olmo-3-7B-Instruct. ALFWorld and WebShop report success rates (%); ScienceWorld reports mean episode score (0–100).
Method
ALFWorld
WebShop
ScienceWorld
ID
OOD
(a) Matched IL continuation
Imitation Learning (IL)
85.71 ± 0.58
79.85 ± 2.79
28.00 ± 1.31
21.94 ± 2.65
IL → IL
84.76 ± 0.89
79.11 ± 3.71
31.40 ± 2.57
22.01 ± 2.01
ACT → IL
89.76 ± 1.88
83.58 ± 3.80
32.00 ± 1.18
32.85 ± 0.15
(b) ACT-stage objective and matched RL controls
Table 3 : Training-objective and continuation controls on Qwen3-8B. ALFWorld and WebShop report success rates (%); ScienceWorld reports mean episode score (0–100). Objective controls use the same action pairs and downstream IL → RL pipeline. Continuation comparisons match six total epochs for IL → IL versus ACT → IL and nine total epochs for IL → RL → RL versus the full ACT pipeline. Bold denotes the highest mean within each block.
Figure 3 : Causal rationale intervention on 400 held-out ALFWorld pairs. (a) Choice accuracy under free generation, direct selection without a rationale, and an injected incorrect rationale. (b) Fraction of examples in which the model follows the action favored by the injected incorrect rationale. Exact values are provided in Table 6 .
Figure 4 : General-reasoning accuracy after training only on ALFWorld agent data. Bars and error bars report the mean ± standard deviation over three runs. Exact values are provided in Table 8 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Base models
Qwen3-4B, Qwen3-8B, Olmo-3-7B-Instruct
Learning rate
2×10−6
Learning-rate scheduler
Cosine
Warmup ratio
0.1
Batch size
64
Group size ( G )
16 (Qwen3-4B); 8 (Qwen3-8B and Olmo 3)
Appendix
Table 4 : Training hyperparameters for Qwen3 and Olmo-3-7B-Instruct.
Benchmark
Domain
Train Pairs
Task Types
Test Samples
ALFWorld
Embodied
10,240
6
140 (ID) / 134 (OOD) episodes
WebShop
Web
3,071
N/A
500 episodes
ScienceWorld
Science
10,240
30
149 episodes
Appendix
Table 5 : Dataset statistics for all training. Train samples are state-action pairs. ID : In-Distribution, OOD : Out-of-Distribution.
Condition
Original Qwen3-8B
ACT checkpoint
Free choice accuracy
45.00
86.50
No-rationale accuracy
46.25
80.75
Incorrect-rationale accuracy
2.25
50.75
Follows injected action ↓
81.75
28.25
Appendix
Table 6 : Exact results for the causal rationale intervention in fig. 3 (%).
Method
ID
OOD
Prompt w/o CoT thinking
13.57
8.96
Prompt w/ CoT thinking
50.71
29.85
ACT stage
71.43
62.69
Imitation Learning (IL)
85.00
83.58
Early Experience
88.57
88.06
ACT → IL
88.57
91.04
Appendix
Table 7 : Cross-size transfer results on Qwen3-4B for ALFWorld. Values are ID and OOD success rates (%). All ACT pairs used by the ACT-based methods were collected from Qwen3-8B.
Method
MATH-500
GPQA-Diamond
Prompt w/o CoT thinking
78.60 ± 0.33
42.93 ± 1.09
Prompt w/ CoT thinking
86.93 ± 0.74
51.52 ± 1.89
Imitation Learning
87.00 ± 0.33
44.61 ± 0.95
Early Experience (Self-Reflection)
86.86 ± 0.25
51.85 ± 0.63
SAND
81.00 ± 0.86
52.53 ± 1.65
Agent-R
86.53 ± 0.81
52.86 ± 1.04
Appendix
Table 8 : Full general-reasoning results corresponding to fig. 4 . ACT denotes the standalone ACT-stage checkpoint, without downstream IL or RL.
Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training signals emerge only within its in-capability region. For tasks where the base policy cannot reach reward states, additional training or external guidance is needed to recover effective learning signals. Rather than relying on costly iterative supervised fine tuning (SFT), we exploit the abundant action data generated in everyday human interactions. We propose \textsc{ActGuide-RL}, which injects action data as plan-style reference guidance, enabling the agentic policy to overcome reachability barriers to reward states. Guided and unguided rollouts are then jointly optimized via mixed-policy training, internalizing the exploration gains back into the unguided policy. Motivated by a theoretical and empirical analysis of the benefit-risk trade-off, we adopt a minimal intervention principle that invokes guidance only as an adaptive fallback, matching task difficulty while minimizing off-policy risk. On search-agent benchmarks, \textsc{ActGuide-RL} substantially improves over zero RL (+10.7 pp on GAIA and +19 pp on XBench with Qwen3-4B), and performs on par with the SFT+RL pipeline without any cold start. This suggests a new paradigm for agentic RL that reduces the reliance on heavy SFT data by using scalable action guidance instead.
Yuxiang Ji, Zengbin Wang, Yong Wang +6
1Xiamen University · 2AMAP, Alibaba Group · Work done during internship at AMAP, Alibaba Group +1
Large language model-based agents make mistakes, yet critique can often guide the same model toward correct behavior. However, when critique is removed, the model may fail again on the same query, indicating that it has not internalized the critique's guidance into its underlying capability. Meanwhile, a frozen critic cannot improve its feedback quality over time, limiting the potential for iterative self-improvement. To address this, we propose learning to internalize self-critique with reinforcement learning(ICRL), a novel framework that jointly trains a solver and a critic from a shared backbone to convert critique-induced success into unassisted solver ability. The critic is rewarded based on the solver's subsequent performance gain, incentivizing actionable feedback. To address the distribution shift between critique-conditioned and critique-free behavior, ICRL introduces a distribution-calibration re-weighting ratio that selectively transfers critique-guided improvements compatible with the solver's own prompt distribution. Additionally, a role-wise group advantage estimation stabilizes joint optimization across the two roles. Together, these mechanisms ensure that the solver learns to improve itself without external critique, rather than becoming dependent on critique-conditioned behavior. We evaluate ICRL on diverse benchmarks spanning agentic and mathematical reasoning tasks, using Qwen3-4B and Qwen3-8B as backbones. Results show consistent improvements, with average gains of 6.4 points over GRPO on agentic tasks, and 7.0 points on mathematical reasoning. Notably, the learned 8B critic is comparable to 32B critics while using substantially fewer tokens. The code is available at https://github.com/brick-pid/ICRL.
Jianbo Lin, Xiaomin Yu, Yi Xin +7
Hong Kong University of Science and Technology (Guangzhou) · Nanjing University · Sun Yat-sen University +3
Language agents can adapt from experience in interactive environments, but current reflection-based methods can only self-correct within a single task instance. Whether such experience can be distilled into reusable lessons that improve performance on future unseen tasks remains unclear. We address this problem by introducing the In-context Training (ICT) task, a framework for evaluating cross-task self-improvement in language agents. In ICT, a reflector model observes trajectories collected by an actor model and generates system prompts intended to improve the actor's performance on future unseen tasks. We then propose an RL-based training pipeline for learning such reflections directly from experience, without human-provided examples. Across ALFWorld and MiniHack, our trained reflectors outperform an untrained baseline on most held-out task families, showing that the ability to learn from experience can itself be learned. In some cases, we observe generalisation beyond the benchmark on which the reflector was trained, to substantially different environments. Finally, we introduce MetaGym, a generic Python library for constructing meta-environments, enabling future research on self-improving language agents.