cs.AIMar 9, 2026

Agentic Critical Training

Authors: Weize Liu, Minghui Liu, Sy-Tuyen Ho, Yongkyun Lee, Andrew Adams Schoen, Souradip Chakraborty, Xiyao Wang, Furong Huang

Organizations: University of Maryland, College Park

Abstract

Imitation learning (IL) teaches language-model agents to reproduce expert actions but not to distinguish them from plausible mistakes. Self-reflection methods expose models to alternatives yet use supervised fine-tuning (SFT) to imitate fixed rationales and actions. We introduce Agentic Critical Training (ACT), which uses reinforcement learning with verifiable rewards (RLVR) to train models to judge actions directly. At each expert-trajectory state, ACT pairs an expert action with an alternative sampled from the initial policy and randomizes their order. The model generates its own reasoning but is rewarded only for selecting the expert action. ACT reuses demonstrations, requires no reference rationales, and allows pair reuse across model sizes. ACT is a warm-up before IL, optionally followed by RL; inference requires no candidate comparison. Across Qwen3-8B and Olmo-3-7B-Instruct on ALFWorld-ID, WebShop, and ScienceWorld, ACT yields average gains of 5.85 points over IL and 4.12 points over IL→\toRL without ACT, while also improving ALFWorld-OOD. With Olmo on ScienceWorld, the full pipeline gains 15.36 points over CoT prompting and 9.23 points over IL→\toRL without ACT. Both ACT→\toIL and the full pipeline outperform supervised reflection baselines. Controls with fixed pairs or matched training durations show that the gains stem from the ACT objective rather than additional data or training. Without reasoning-specific post-training, the standalone ACT checkpoint achieves the highest mean among evaluated models on MATH-500 and GPQA-Diamond, showing that action comparison complements generation.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Learning Agentic Policy from Action Guidance

    May 12, 2026Yuxiang Ji, Zengbin Wang, Yong Wang +6Agentic RLReinforcement Learning

  2. ICRL: Learning to Internalize Self-Critique with Reinforcement Learning

    May 13, 2026Jianbo Lin, Xiaomin Yu, Yi Xin +7RL for Language ModelsLLM Self-Refinement

  3. Training Language Agents to Learn from Experience

    May 19, 2026Yuval Shalev, Zifeng Ding, Mateja JamnikContinual Learning for LLM AgentsSelf-Improving Agents