Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rates across four base models of different scales and architectures, including three not used during critic training. On the MuSiQue benchmark, the trained critic improves Qwen3-4B's success rate by more than 25 percentage points, surpassing the performance of Kimi K3 without a critic. The same critic also yields gains on out-of-domain interactive benchmarks, including τ3 and DeepDive, with no additional training. Our results show that agents can decide when to seek help from a critic at inference time and that outcome-based critic training can produce guidance that transfers across base models and task domains.
Figures & tables
Figure 1: Training and using Caddie . (a) Training: the frozen agent πb rolls out a task and a step k is chosen randomly; at the prefix sk the critic πcθ samples G critiques; each critique is then appended to the prefix and the frozen agent rolls out N continuations until the end of the episode. The terminal rewards are normalized within the group and update only the critic with DAPO. (b) Inference: the agent decides when to consult the critic through the call_critic tool and receives the critique as the tool output.
MuSiQue
Critic
Attached
Wiki
ALFWorld
No critic
21.56±2.07
10.38±1.59
23.79±3.01
Prompted Qwen3-4B
23.25±2.37
16.06±2.19
20.24±2.89
Prompted Kimi K3
29.13±2.46
19.19±2.23
30.22±2.80
Caddie (ours)
46.88±2.68
21.13±2.38
49.72±3.42
Table 1: In-domain success rates (%) with Qwen3-4B as the base model. Each Caddie critic is trained in the corresponding task setting. Entries report mean ± SEM across tasks, with eight rollouts per task. Bold marks the highest mean; italics mark other means whose ± SEM intervals overlap the leader’s (not a significance test).
Prompted critic
Base model
No critic
Qwen3-4B
Kimi K3
Caddie (ours)
Attached passages
Qwen3-30B-A3B
28.44±2.45
34.75±2.82
36.00±2.86
41.63±2.83
Qwen3.8-27B
46.63±2.93
45.31±2.84
41.13±2.77
55.94±2.93
Kimi K3
38.56±2.87
34.44±2.63
30.19±2.36
51.00±2.76
Wikipedia retrieval
Table 2: Transfer across base models on MuSiQue. Success rates (%) are reported as mean ± SEM across tasks, with eight rollouts per task. Caddie critics are trained with Qwen3-4B and evaluated with each base model without further adaptation. Bold marks the highest mean; italics mark other means whose ± SEM intervals overlap the leader’s (not a significance test).
τ3
Critic
DeepDive
Retail
Airline
Base model: Qwen3-4B
No critic
26.95±2.38
44.69±1.73
31.25±1.57
Prompted Qwen3-4B critic
28.91±1.39
42.50±0.98
31.88±2.49
Prompted Kimi K3 critic
27.34±0.72
40.31±1.53
27.50±2.83
Agent-RRM
29.30±0.75
40.63±2.15
33.75±2.63
Table 3: Transfer of MuSiQue-trained critics to unseen tasks. Entries report success rates (%): the mean over all tasks and seeds ± the SEM across the eight seed-level aggregate accuracies. Shading marks our critics. Bold marks the highest mean; italics mark other means whose ± SEM intervals overlap the leader’s (not a significance test).
Figure 2: Solve rates on tasks where Caddie delivered feedback. Navy: feedback-bearing rollouts; lavender: no-critic rollouts on the same tasks. This generator-selected subset gives a descriptive comparison, not a causal estimate. Q4/Q38 denote Qwen3-4B/Qwen3.8-27B; DD denotes DeepDive.
Figure 3: Matched τ3 Retail trajectories after the same change in user intent. Each path summarizes the critic advice and the frozen-generator continuation.
Figure 4: DeepDive solve rate by advice type. Lavender (retained baseline): no-critic rollouts on the tasks where that advice type was given; navy: rollouts that received it.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: MuSiQue validation success on a 20-task subset for four reward variants, shown up to 2,000 training steps. All curves use an exponential moving average with weight 0.5 on the current evaluation. Confidence bands are unavailable; these smoothed traces illustrate training dynamics but do not establish differences in performance, convergence speed, or reward variance.
Table 4: Optimization settings for MuSiQue-attached, MuSiQue-wiki, and ALFWorld.
Critic
Attached
Wiki
Qwen3-4B before GRPO
No critic
21.56±2.07
10.38±1.59
Qwen3-4B after GRPO
No critic
49.69±2.98
19.69±2.36
Caddie trained with original generator
51.69±2.87
24.44±2.64
Caddie trained with GRPO generator
63.56±2.98
23.38±2.63
Appendix
Table 5: Critic feedback on the GRPO-trained Qwen3-4B generator (step 1,120), with the original instruction-tuned generator as a reference. All evaluations with the GRPO-trained generator require at least three searches. Success rates (%) are mean ± SEM across 200 tasks, with eight rollouts per task. Both critic rows use self-call with at most three critiques per trajectory. Bold marks the highest mean; italics mark other means whose ± SEM intervals overlap the leader’s (not a significance test).
Critic training
Inference protocol
Retrieval
Base model
k
Fixed step
Self-call
Evaluation: Qwen3-4B, attached
Wiki
Qwen3-4B
4
28.00±2.47
26.88±2.46
Attached
Qwen3.8-27B
5
30.50±2.32
28.88±2.60
Evaluation: Qwen3.8-27B, wiki
Wiki
Qwen3-4B
4
25.94±2.63
19.06±2.13
Appendix
Table 6: Selected MuSiQue comparisons where fixed-step intervention has a higher mean than self-call. Both critics are Qwen3-4B models; their training retrieval settings and base models are listed below. Success rates (%) are mean ± SEM across 200 tasks, with eight rollouts per task. Self-call permits at most one critique. Bold marks the highest mean; italics mark other means whose ± SEM intervals overlap the leader’s.
Base model
k=2
k=3
k=4
k=5
Terminal-outcome critic, N=1
Qwen3-4B
22.12±1.92
17.75±1.66
30.88±2.18
33.69±2.35
Qwen3-30B-A3B
36.44±2.42
35.44±2.19
41.50±2.46
36.38±2.29
Terminal-outcome critic, N=5
Qwen3-4B
23.06±1.96
13.81±1.37
27.88±2.14
32.56±2.30
Qwen3-30B-A3B
34.38±2.36
25.19±1.88
35.94±2.17
34.75±2.32
Appendix
Table 7: Sensitivity to neighboring intervention steps on MuSiQue-attached. Each row keeps the base model and critic checkpoint fixed. Entries are success rates (%; mean ± SEM across 200 tasks, with eight rollouts per task). Bold marks the highest mean; italics mark other means whose ± SEM intervals overlap the leader’s.
Figure 6: Critic requests per called rollout.
Figure 7: DeepDive accuracy by the number of critic requests.
Figure 8: How generators use critic feedback and how the advice classes vary across domains.
Figure 9: τ3 solve rate by advice type.
Figure 10: Qwen3.8-27B on DeepDive with the MuSiQue-wiki Caddie critic. Left: paired solve-rate change (percentage points) on the tasks where the critic was called, by task family. Right: share of called rollouts that improved or degraded after the generator either committed to an answer or performed another search.
Figure 11: Qwen3.8-27B solve rate by τ3 task family. Each group uses the same tasks and replicates for every critic; A/W denote the MuSiQue-attached/wiki Caddie checkpoints.
Large language model (LLM) agents are increasingly deployed in long-horizon, interactive, and stateful environments. In these settings, a single wrong action, such as refunding the wrong purchase, can cause irreversible task failure and must be intercepted before execution. Such failures may not appear in every single run, but can emerge across repeated trials, making reliability across steps and trials critical. However, ensuring agentic reliability is challenging: even frontier LLMs struggle to explain why an action may be wrong, especially in long, intertwined trajectories governed by domain-specific policies. Much recent work relies on prompt-based critique agents, while optimization-based methods lack a systematic way to produce rich verification rationales for training. We address this gap with CAST, a critique-aware training framework that converts sparse task outcomes into action-level supervision for critique learning and policy optimization. CAST analyzes agent trajectories to synthesize structured rationales explaining action validity under partial observability. The resulting critique model is used to construct critique-aware training data for optimizing the policy model. Fine-tuning Qwen3-family models on dynamic tool-calling benchmarks, CAST improves reliability across domains, outperforming GPT-OSS-120B by over 10% pass^4 on Retail tasks and yielding an additional 9% improvement on Telehealth in an out-of-domain setting. These results demonstrate that critique-aware training improves the robustness of LLM agents in realistic dynamic environments.
LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong step can waste interaction budget or trigger irreversible side effects long before the final failure is observed. Reliable deployment therefore requires \emph{step-level confidence estimation}: a calibrated probability that each proposed action is productive, available \emph{before} the action is executed. Existing LLM confidence estimators are designed to score a response from the given prompt, but agent confidence also depends on execution consequences: whether similar actions in similar situations actually advanced the task after the environment responded. We introduce the \method (\methodshort), a self-evolving critic framework in which an LLM critic accumulates evidence from its own past judgments and their observed consequences. After each trajectory, a hindsight LLM that sees the full execution feedback votes on whether each step was productive. The resulting pseudo-labels populate a memory bank from which related productive and unproductive experiences are retrieved into the critic's prompt whenever a similar step recurs. \methodshort requires no training and uses no ground truth step labels. Across three agent benchmarks and three critic backbones, \methodshort attains the best calibration (ECE and Brier) and ranking (AUC) in every dataset--critic combination, reducing ECE by up to 54% relative to the strongest training-free baseline.
Yaopei Zeng, Congchao Wang, JianHang Chen +3
Pennsylvania State University · Virginia Tech · Purdue University +1
Chain-of-thought (CoT) reasoning is widely used in language-model agents, but prior work has shown that verbalized CoT is not always faithful and may instead reflect post-hoc reasoning, which means the model already knows the answer before reasoning. We therefore ask what CoT training is actually improving: is the model getting better at changing its action through generated reasoning, or is it getting better at predicting the action directly from the prompt? We study this question by comparing \emph{prompt actions} (predicting action without CoT) with CoT actions (predicting action with CoT). Across checkpoints, prompt-action quality improves substantially. While interacting with the environment, the relative advantage of CoT actions over prompt actions remains similar, showing that CoT training does not widen the advantage of CoT reasoning, and it helps to improve the quality of prompt actions. We further find that later checkpoints are less likely to revise the action in response to CoT, suggesting greater reliance on the prompt. Motivated by these patterns, we selectively mask action-token supervision on a fraction of training examples. This intervention improves out-of-domain generalization.
Jingyu Liu, Zhiwen Wang, Yuxin Jing +2
Gaoling School of Artificial Intelligence Renmin University of China, Beijing, China · ByteDance · Beijing Key Laboratory of Research on Large Models and Intelligent Governance +1