My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning
Authors: Yihua Zhu, Qianying Liu, Weixu Qiao, Xuan Ren, Weiwei Xu, Wenbo Li, Wei Wang, Ruijia Chen, +5 more
Organizations: Alibaba Group · Kyoto University · NII LLMC · Peking University · University of California, Los Angeles · Arizona State University · The Chinese University of Hong Kong, Shenzhen · Tsinghua University · University of the Chinese Academy of Sciences
Agentic reinforcement learning (RL) has emerged as a powerful approach for training large language model agents on multi-step tasks, yet reliance on terminal outcome rewards creates two credit-assignment problems, particularly in long-horizon tasks. First, same-outcome rollout groups provide no learning signal from terminal rewards. Second, terminal rewards provide only trajectory-wide feedback, making it difficult to identify which decisions caused a failure. Recent work supplements terminal rewards with finer-grained information from trajectory analysis, such as natural-language reflections on intermediate decisions and errors. However, natural-language diagnoses are difficult to use directly for credit assignment: their error claims may be unreliable, and they do not quantify how much each error should affect learning. We propose Self-Diagnosis-guided Terminal Credit Redistribution (FAULT), which turns diagnosed errors into explicit step-level credit anchored by terminal outcomes. FAULT checks diagnostic evidence and learns relative error costs from task outcomes. During training, the policy and self-diagnoser co-evolve, while error costs are updated online from recent outcomes. On ALFWorld, FAULT recovers learning signals from same-outcome groups, reaching 95% signal coverage versus 41% for GRPO and 72% for GiGPO, while better localizing credit to specific error steps. Across two model scales, FAULT delivers strong. improvements on the long-horizon ALFWorld and WebShop tasks while remaining competitive on short-horizon Search-based QA.
Figures & tables
Figure 1 : Signal coverage, step localization and final performance. (a) Signal coverage. On ALFWorld with shared diagnosis-SFT initialization, hatched bars show the group mix and solid bars the groups with usable signal. FAULT uses 95% of groups, including every all-fail group and most all-success groups. GRPO uses only mixed groups ( 41% ); GiGPO also uses some all-success groups but no all-fail groups ( 72% ). Thus, FAULT can learn from differences in intermediate behavior even when terminal rewards are identical, making both successful and failed rollouts useful beyond their final outcomes. (b) Localization. Mean reciprocal rank of the decisive error step a blind LLM judge marks in each failed trajectory, with steps ranked by each method’s penalty. FAULT ( 0.494 ) nearly doubles random ( 0.266 ) while GiGPO ( 0.303 , same-configuration rerun) stays close to it, so FAULT’s step-level credit is far more accurate. The gains highlight the value of diagnostic evidence in identifying which decisions within a failed trajectory need correction and directing step-level feedback to those decisions. (c) Outcome. Qwen3-4B results from Table 1 . FAULT leads in ALFWorld success rate and WebShop task score on these longer-horizon tasks, remains competitive with SEED on shorter-horizon Search-based QA, and outperforms GRPO and GiGPO initialized from the raw backbone on all three.
Figure 2 : Overview of FAULT. Stage I initializes a shared actor and diagnoser by SFT on teacher diagnoses and learns error prices from terminal outcomes. In Stage II , we diagnose errors in on-policy rollouts, price them from terminal outcomes, and redistribute a bounded penalty budget across diagnosed steps for policy updates. The actor, diagnoser and prices co-evolve through shared model updates and online price refits.
ALFWorld
Search-based QA
WebShop
Method
Pick
Look
Clean
Heat
Cool
Pick2
All
NQ
Triv
Pop
Hotp
2Wk
MuS
Bam
Avg
Score
Succ.
Qwen3-1.7B
Vanilla
8.3
22.2
6.5
0.0
0.0
0.0
6.0
29.7
48.2
35.0
23.7
19.6
6.4
17.6
25.7
51.0
2.3
GRPO
66.7
38.9
35.5
43.5
52.4
17.6
43.3
39.3
57.1
43.3
35.7
33.9
11.7
28.0
35.6
66.5
39.1
GiGPO
58.3
27.8
29.0
65.2
38.1
29.4
41.8
39.1
56.6
44.4
33.2
29.2
11.2
23.2
33.8
78.8
49.2
SEED
79.2
61.1
71.0
65.2
61.9
47.1
65.7
42.3
58.5
44.8
38.4
36.9
15.4
36.0
38.9
86.8
73.4
Table 1 : Performance on ALFWorld, Search-based QA, and WebShop. ALFWorld reports success rate (%) by task family and over all 134 unseen tasks (All); Search-based QA reports exact-match accuracy (%) by dataset and its unweighted average (Avg); WebShop reports mean task score and strict purchase success (%) over 256 episodes. Rows marked (diag-SFT) start RL from our Stage 1 cold-start model. Within each backbone, bold and underlined mark the highest and second-highest reported values per column (ties share formatting).
Figure 3 : Relative signal strength during training. Stacked areas show the contributions of mixed, all-success, and all-fail groups to the signal-retention Index. Black lines show totals for 20 -update windows; dotted lines show run-pooled values. Both baselines start from diagnosis-SFT.
Figure 4 : Placement and concentration of step-level credit. (a) Offset of the most penalized step from the judge’s decisive step over 300 failed trajectories: FAULT in dark bars, GiGPO (same-configuration rerun) in light bars. (b) Cumulative distribution of the normalized credit entropy on same-outcome groups; curves further left are more concentrated, and the legend gives the mean effective number of credited steps. Baselines start from the diagnosis-SFT checkpoint.
Figure 5 : ALFWorld training trajectory length. (a) Mean active steps per training update (thin lines) and five-update moving averages (thick lines); shading marks updates 151 – 160 . (b) Means over updates 151 – 160 . All runs use a 30 -step limit but differ in initialization and training recipe.
ALFWorld
Variant
Pick
Look
Clean
Heat
Cool
Pick2
All
FAULT
91.7
100.0
93.5
87.0
85.7
88.2
91.0
w/o online pricing
79.2
100.0
83.9
95.7
90.5
94.1
89.6
Cosine λ→ constant λ
70.8
100.0
77.4
73.9
85.7
70.6
79.1
Shared diagnoser → fixed judge
91.7
94.4
77.4
78.3
81.0
76.5
82.8
GRPO + diagnostic penalty
79.2
100.0
71.0
87.0
85.7
41.2
77.6
Table 2 : Ablations and design variants on ALFWorld with Qwen3-4B-Instruct-2507: success rate (%) on the 134 unseen tasks, evaluated as in Table 1 .
Figure 6 : Diagnostic quality during RL. Concrete-L2 precision, recall, micro-F1 and macro-F1 against the fixed teacher reference. Shading marks the late λ→0 window (updates 100 – 160 ). Single seed; training rollouts.
Symbol
Meaning
Where
Indices, sets and constants
q∼D
task and its rollout group; D : training task distribution
§ 2.1
i,j∈{1,…,G}
rollout indices in a group; G rollouts per task
§ 2.1
t,r∈{1,…,Tq,i}
active steps; Tq,i : steps executed, not the step limit
§ 2.1
k∈{1,…,K}
priced error class; C={c1,…,cK} : frozen priced set
§ 2.2
u∈{1,…,U} ; v
actor update (batch), U in total; accepted price refit
§ 2.4
Table 3 : Notation. “Where” gives the place where the symbols are first introduced.
Setting
ALFWorld
WebShop
Search-based QA
1.7B
4B
1.7B
4B
1.7B
4B
Supervised initialization
SFT learning rate
10−5
SFT learning-rate schedule
10% warm-up; cosine decay
SFT epochs / global batch
2 / 32
7 / 176
2 / 32
Policy optimization and rollout generation
Table 4 : Key training and method hyper-parameters for the six FAULT runs. Shared values span model or benchmark columns.
Figure 8 : Usable signal over training. Stacked shares of all groups with a usable range at τ=0.05 , measured in 20 -update windows at their final update. The dark line is total coverage; titles report run-wide values. Same-outcome groups (red and gold) dominate and grow, with FAULT retaining broad coverage of them. FAULT uses scores S and the baselines use optimizer advantages. Single-seed runs; both baselines start from diagnosis SFT.
Figure 9 : Threshold sensitivity and pooled signal strength. (a) Coverage as τ increases, using each method’s mixed-group mean; the dotted line marks τ=0.05 . (b) Pooled Index split by group type; the mixed contribution equals its group share. (c) Same-outcome contributions on a finer scale. Coverage uses S for FAULT, whereas its Index uses P ; baselines use advantages for both. These measure contrast, not independently verified correctness.
Figure 10 : Diagnostic penalties on successful trajectories. Each point summarizes 20 updates. (a) Verified and L1-majority report rates among valid diagnoses, and charge rate ( P>0.001 ) among all successes. (b) Median budget P and interquartile range for successes and failures. (c) Fraction of the 2,560 trajectories per window reaching the cap; labels give counts. Windows are placed at their final update. Diagnoser and prices evolve during training; single seed.
Figure 11 : Additional localization and concentration results. (a) Hit@ k within one step of the blind judge’s label, using the trajectories of Figure 4 (a) and 48 GRPO trajectories; the dashed curve is random ranking. (b) Mean ENS on all-fail groups in 20 -update windows, placed at their final update. Both baselines start from diagnosis SFT.
Figure 12 : Negative-credit allocation to wasted steps. (a) Mean targeting ratio for invalid actions and format-valid no-ops under shared mechanical labels; 1× is uniform allocation. (b) No-op targeting over training. Credit is δt for FAULT and max(−At,0) for the baselines. All analysed failures reach the step limit. Baselines: GRPO (diag-SFT) and GiGPO (diag-SFT); offline analysis of single-seed runs.
ALFWorld
WebShop
Cold-start pool
Diag. F1
Non-zero prices
SFT SR
RL SR (draws)
Score
Succ.
1,440
0.617
6/20
24.6
74.6
73.8
60.2
2,880
0.648
8/20
31.3
82.1
85.6
67.6
4,320
0.668
3/20
30.6
83.6
82.3
70.3
Table 5 : Cold-start data scale with Qwen3-4B-Instruct-2507. ALFWorld reports named-class diagnosis F1 of the SFT model, non-zero initial prices, pre-RL SFT success rate and success rate after 160 RL updates. WebShop reports post-RL task score and strict success rate. Evaluation protocols are specified above.
Updates
GRPO raw-init
SEED
FAULT
w/o online pricing
156–160
21.89
14.54
11.43
12.03
151–160
21.69
13.77
12.03
11.38
141–160
22.09
14.34
12.56
11.70
121–160
22.21
14.38
12.64
12.00
Table 6 : Mean training trajectory length over four windows ending at update 160 , the last update shared by all runs. Each value equally averages the per-update mean number of active interactions. All runs use a 30 -step horizon; these are single-run training summaries, not test-time estimates or confidence intervals.
Figure 13 : Penalty allocation and behavior frequencies during training. Points summarize 20 -update windows and are placed at the final update. (a) Share of FAULT’s diagnostic penalty mass assigned to each mechanical step class. (b) Format-valid no-op rate. (c) Invalid or missing action rate. Rates use all recorded steps as the denominator; the remaining-step category in (a), labeled effectful, includes steps without subsequent feedback. GRPO starts from diagnosis SFT; SEED is a reference with a different training mechanism. GiGPO is omitted because its original logs lack rollout text. Single-seed observational comparisons.
Pricing class
βk
σk=\sd(x⋅,k)
wk
wk
E7/_other (non-executable output, other)
−0.10791
0.017713
1.9115×10−3
1.0000
E1.state_requirement_misread
−0.07593
0.015166
1.1515×10−3
0.6024
E2.failure_cause_misattribution
−0.06719
0.013516
9.0816×10−4
0.4751
E6/_other (premature abandonment, other)
−0.04978
0.012275
6.1109×10−4
0.3197
E3.redundant_state_toggle_repeat
−0.02184
0.010325
2.2551×10−4
0.1180
E1.state_history_loss
−0.01125
0.008723
9.8129×10−5
0.0513
Table 7 : Cold-start prices on ALFWorld with Qwen3-4B-Instruct-2507. wk=max(−βk,0)σk and wk=wk/maxjwj (Eq. ( 17 )). The remaining 12 pricing classes have βk>0 and wk=0 .
Figure 14 : Coefficient dynamics for three illustrative error classes. Sign-gated moving averages at 14 refits per run; panels use different vertical scales. GRPO and GiGPO (diag-SFT) track coefficients without applying them, the frozen-price variant retains its initial prices, and the full method deploys updated prices. E7 residual collects unnamed errors within the non-executable-output family.
Figure 15 : Changes in price rankings and relative price allocation. (a) Spearman correlation between cold-start prices and each run’s online weights. Only the full method deploys the updated weights; the other curves are monitoring estimates. (b) Search-state amnesia under the full method: its share of total deployed class weight (solid) and its occurrence rate among tracker-window trajectories (dashed). Shading marks updates 30 – 60 . The two series have Pearson correlation +0.996 over updates 70 – 160 ; the class’s own normalized weight remains 1 . Each curve represents 14 refits from one run.
Figure 16 : Coefficient dynamics for all 20 pricing classes. Sign-gated moving averages at 14 refits for five runs, with a separate vertical scale per panel. Coincident curves may be hidden by the full-method curve drawn last; several runs remain at zero in panels 3, 5 and 9. Residual classes collect errors not assigned to a named L2 class within an L1 family. Curves show estimator coefficients, not deployed prices.
Figure 17 : Diagnostic penalties within an all-fail group. (a) Budgets for eight rollouts, all failing at 30 steps. (b) Allocation for the selected trajectory: approximately 82% at the judge’s decisive step and 18% at a later revisit (Table 8 ). The group contributes to Figure 8 ; the case is selected for illustration.
Step
Action
Environment feedback (abridged)
δt
0–1
go to / open fridge 1
fridge opened: a cup and a potato, no lettuce
0
2–7
first sweep
countertop 1; cabinets 2, 3 (opened, both empty); sinkbasin 1: four further locations, each visited for the first time
0
8
go to fridge 1
first revisit of a searched location; contents unchanged since step 1
2.60 (82%)
9–13
search starts to cycle
countertop 1 and cabinet 2 revisited; cabinet 4 and diningtable 1 new
0
14
go to cabinet 3
revisit of a cabinet known to be empty since step 6
0.58 (18%)
15–28
cycling continues
nine more steps on already-searched locations, four on new ones (diningtable 2, cabinet 1, sidetable 1), one inventory check (“not carrying anything”)
0
Table 8 : Walkthrough of the selected trajectory in Figure 17 (b), abridged from the raw episode. Values of δt are rounded shares of the diagnostic budget P=3.18 ; the two largest allocations fall on revisits of searched locations.
Figure 18 : Transcripts and step penalties. (a) GRPO (diag-SFT) and (b) FAULT failures from separate all-fail groups, each with its per-step credit: GRPO’s negative optimizer advantage max(−At,0) and FAULT’s diagnostic allocation δt . The blind judge selects steps 13 and 8 , respectively (shaded); GRPO marks only three unparseable steps, while FAULT allocates approximately 82% to the judged step. (c) Negative advantage for a GiGPO (diag-SFT) all-fail trajectory, marking only format violations; no transcript was stored. Action and feedback text is verbatim; italic notes are ours. Task instances differ.
Figure 19 : Recovery within a successful trajectory, steps 0–9. The agent makes three ineffective heating attempts. Charge tags show the step allocation of budget P=5.52 ; the case continues in Figure 20 .
Figure 20 : Recovery within a successful trajectory, steps 10–18. Continued from Figure 19 . Examination reveals the tomato is still cold, after which the agent corrects its belief and completes the task. Most of the budget targets ineffective heating attempts, although the corrective examination also receives a penalty.
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work introduces privileged self-distillation for credit assignment, providing denser supervision, but it remains unclear how such local signals should represent sequential credit. We propose AgentOPSD, a critic-free, recursive method for turn-level credit assignment in agentic reinforcement learning. AgentOPSD aggregates token-level teacher-student log-probability gaps into turn-level evidence and recursively updates a Bayesian belief state in log-odds space. This yields a principled reweighting scheme that converts sparse outcome supervision into turn-level credit signals and identifies pivotal turns through the marginal belief revision between consecutive states. The method is fully compatible with standard policy optimization and requires neither an additional critic nor extra rollouts. We evaluate AgentOPSD on ALFWorld, WebShop, and Search-QA using Qwen2.5 models at two scales (3B and 7B). AgentOPSD outperforms GRPO and strong self-distillation baselines, achieving 89.1% success on ALFWorld with Qwen2.5-7B. Ablation studies attribute the gains to turn-level aggregation and history-dependent recursive belief updates.
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10
Tsinghua University · Meituan · Zhejiang University
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advantage over all action tokens. This outcome signal is useful but structurally incomplete: it punishes useful exploration in failed rollouts and reinforces redundant or regressive actions in successful rollouts. We propose TRIAGE, a role-typed credit assignment framework that adds a semantic role axis to outcome credit. A structured judge classifies each segment as decisive progress, useful exploration, no-progress infrastructure, or regression, and a fixed role-conditioned rule maps these labels to bounded segment-level process rewards. This keeps verifier outcomes as the source of optimization direction while correcting the two main blind spots of outcome-only credit. We further show that role-conditioned credit is the optimal segment-level correction expressible from role labels alone -- a projection of the per-segment advantage residual onto the role variable -- so that the fixed role constants reduce advantage estimation error whenever the judge is reliable, and we connect this to lower-variance policy gradients. Across ALFWorld, Search-QA, and WebShop, TRIAGE improves success rates over GRPO for two policy models and outperforms both a scalar judge-derived process reward and an outcome-supervised shared-backbone value baseline. Ablations show that the gain comes from role typing rather than merely adding dense rewards: reliable detection of regression inside successful trajectories is the dominant contributor, while exploration credit provides a consistent secondary gain; on completed ALFWorld and WebShop rollouts, TRIAGE also reduces environment-facing turns by an additional 10.4% and 14.8% relative to GRPO.
Yuanda Xu, Zhengze Zhou, Hejian Sang +6
1LinkedIn Corporation · 2Harvard University · 3Johns Hopkins University +1
Contemporary reinforcement learning with verifiable reward methods post-train language models on multi-step reasoning by assigning a single outcome reward uniformly across all tokens in a trajectory. Such uniform assignment ignores which steps contributed to success or failure. Improving credit assignment can address this limitation by enabling targeted refinement of faulty reasoning steps, rather than updating entire trajectories uniformly. Resets are one such simple mechanism, enabling more precise credit assignment by returning to an intermediate state and resampling counterfactual continuations, so that outcome differences can be attributed to decisions made at that point. We propose two such methods: Random-Reset Policy Optimization (RRPO), where reset states are drawn randomly from reasoning steps, and Self-Reset Policy Optimization (SRPO), where the model self-localizes the erroneous step in an incorrect trajectory and resets there. We analyze these methods within the Conservative Policy Iteration (CPI) framework. Extending CPI with a credit-assignment oracle that targets improvable states yields provable improvements over random resets. Across models and reasoning benchmarks, SRPO consistently outperforms standard GRPO and RRPO by sampling multiple suffix continuations at a self-localized reset and learning from their rewards, using only the model itself with no external supervision.
Ankur Samanta, Akshayaa Magesh, Ayush Jain +7
1Meta AI · 2Columbia University · *Work done at Meta +2