Organizations: School of Data Science, Fudan University · Shanghai Innovation Institute · Beihang University · Simple AI · Tsinghua University · Tuojing Intelligence · The University of Hong Kong
Outcome-based reinforcement learning (RL) post-trains vision--language--action policies using terminal success signals, but assigns the same trajectory-level advantage to every action chunk. A failed episode can thus penalize useful early actions as if they caused the failure. Existing approaches seek finer-grained feedback through learned evaluators, adding task-specific supervision or additional model training. We explore, for the first time to our knowledge, whether physical relations across trajectories can provide action-chunk credit in embodied RL from terminal outcomes alone, without an auxiliary evaluator. The key insight is that rollouts reaching corresponding physical situations can serve as references for one another: their terminal outcomes provide evidence for assessing local progress. We introduce Physical Relations for Inferring Credit from Episodes(PRICE), with two components: (i) a physical relational graph that pools current and historical outcomes at corresponding chunk boundaries to estimate success potentials; and (ii) confidence-gated credit assignment that uses changes in these potentials to refine trajectory-level supervision. Our analysis connects oracle potential changes to the terminal-success objective and provides a finite-sample directional bound for outcome-independent evidence pools. Independent continuation tests show that PRICE's retained credits align with local progress, while experiments on LIBERO, RoboTwin 2.0, and real robots demonstrate improved task success over outcome-based baselines and faster learning.
Figures & tables
Figure 1: Motivation for PRICE. PRICE derives local action-chunk credit by comparing physically corresponding states across rollouts, using only terminal outcomes and no auxiliary evaluator.
Figure 2: Overview of PRICE. (a) Visual–proprioceptive descriptors match physically corresponding states across rollouts. (b) Cross-trajectory terminal outcomes determine node success potentials. (c) Confidence-gated endpoint differences provide action-chunk credit that augments the GRPO advantage for policy optimization.
Method
Policy input
Spatial
Object
Goal
Long
Avg.
VLA baseline
OpenVLA ( Kim et al., 2024 )
T+I
84.7
88.4
79.2
53.7
76.5
OpenVLA ∗ -Full ( Fei et al., 2026 )
T+I
91.6
95.3
90.6
86.5
91.0
Evaluator-based RL
TGRPO ( Chen et al., 2025b )
T+I
90.4
92.2
81.0
59.2
80.7
GRAPE ( Zhang et al., 2024 )
T+I
88.5
92.1
83.1
57.2
80.2
Table 1: Performance comparison on the LIBERO benchmark. We report task success rates (%). Rows below the dashed line are our results; other values follow the unified comparison of Fei et al. (2026) . T, W, P, and I denote third-person images, wrist images, proprioception, and language instructions, respectively. Full and One denote full-shot and one-trajectory-per-task SFT, respectively.
Method
Handover Block
Lift Pot
Move Can Pot
Pick Dual Bottles
Place Container Plate
Place Empty Cup
Avg.
SimpleVLA-RL ( Li et al., 2026a )
57.8
64.1
61.2
68.3
82.1
94.2
71.3
RLinf-VLA ( Zang et al., 2025 )
70.31
70.31
83.59
92.96
95.31
94.53
84.5
Feat2Go ( Shu et al., 2026 )
76.6
78.9
91.4
93.8
96.9
95.3
88.8
OpenVLA-OFT (SFT)
28.1
3.1
9.4
20.3
54.7
75.8
31.9
+ PRICE (Ours)
79.7
80.5
90.6
94.5
97.7
96.9
90.0
↑ 51.6
↑ 77.4
↑ 81.2
↑ 74.2
↑ 43.0
↑ 21.1
↑ 58.1
Table 2: Performance comparison on six selected RoboTwin 2.0 manipulation tasks. We report task success rates (%) and the average across the six tasks. The OpenVLA-OFT (SFT) row is the initialization used for PRICE; results for the other compared methods are from the cited works.
Figure 3: Training efficiency on LIBERO. (a–b) Rollout success during training on Goal and Spatial. Solid curves show seven-step moving averages and faint lines show raw observations; dashed guides and x-axis ticks mark the steps for 90% and 95% success. (c) Training steps of SRPO and PRICE across four suites. SRPO counts are from Fei et al. (2026) ; corresponding success rates are reported in Table 1 .
Figure 4: Real-robot performance and representative executions. (a) Success rates of the initialization, GRPO-AWR, and PRICE-AWR on Remote Insertion and Produce Sorting under the HiFi-UMI evaluation protocol. (b) Chronological keyframes show remote transport, alignment, and insertion into the storage box (top), and fruit grasping, transport, and release into its designated tray during Produce Sorting (bottom).
Variant
History
Gate
Spatial
Object
Avg.
GRPO
×
×
93.2
94.6
93.9
PRICE w/o Historical Evidence
×
✓
93.4
94.6
94.0
PRICE w/o Gate
✓
×
95.8
97.9
96.9
Full PRICE
✓
✓
97.0
99.6
98.3
Table 3: Component ablations on LIBERO-Spatial and LIBERO-Object. All variants use 63 training steps with the same initialization, interaction budget, and optimization settings within each suite. We report task success (%) and the average across the two suites.
Figure 5: Credit analysis on LIBERO-Spatial. Rows denote trajectory windows and columns action chunks. (a) GRPO advantage sign. (b) Gated PRICE credit from (e). (c) Independent MC progress over the pooled evidence in (e). (d–e) Pre-gate contrasts without/with history at fixed nodes. (f) MC-based credit evaluation for all nonzero candidates in (e) and their rejected/retained subsets.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Setting
Parameter
Setting
Simulation training and GRPO
PRICE in simulation
Optimizer
AdamW
Credit weight λ
0.2
Learning rate
5×10−6
Matching threshold η
0.93
Learning-rate schedule
Constant
Gate parameter δedge
0.15
Cases per raw rollout batch
64
Summaries per node Chist
4
Rollouts per case
8
Nodes per task Cnode
1,024
Appendix
Table 4: Core training hyperparameters. Optimization batch sizes count trajectories and are global across eight GPUs. Simulation and real-robot credit weights are reported separately.
LIBERO
Method
Spatial
Object
Goal
Long
Avg.
PRICE
99.2±0.2
99.6±0.2
99.4±0.3
97.6±0.5
99.0±0.2
Appendix
Table 5: Main PRICE results (%, mean ± SD) across three independent training runs.
Suite
Success (%)
SimpleVLA-RL
PRICE
Reduction (%)
Goal
90
441.9
272.8
38.3
95
625.2
517.2
17.3
Spatial
90
329.0
221.2
32.8
95
714.6
544.3
23.8
Appendix
Table 6: Cumulative GPU hours at the marked milestones in Figure 3 (a–b). Reductions are relative to SimpleVLA-RL.
Training round
DirAcc (%) ↑
RankCorr ↑
AlignedGap (pp) ↑
16
95.8
0.67
+24.6
32
95.8
0.71
+27.8
63
94.3
0.68
+26.2
Appendix
Table 7: Online credit alignment on LIBERO-Spatial. Of 1,200 audited chunks per stage across three PRICE runs, 24, 48, and 36 receive nonzero credit at rounds 16, 32, and 63, respectively.
Selector
DirAcc (%) ↑
AlignedGap (pp) ↑
Confidence gate
97.5
+28.5
Endpoint-count selection
65.9
+4.6
Random selection
77.2
+12.8
Appendix
Table 8: Credit selection at equal coverage on LIBERO-Spatial. Each selector retains 45 of the same 123 candidates. MC ties are excluded from DirAcc; the random row reports means over 100,000 subsets.
Figure 6: Credit selection and local progress on LIBERO-Spatial. (a) Raw PRICE credit versus MC progress for the same 123 candidates as Figure 5 ; blue and gray denote the 45 retained and 78 rejected candidates, respectively. Shaded quadrants indicate agreement in sign. (b) Per-chunk aligned MC gaps ge (pp), whose means give AlignedGap. Vertical jitter separates points; thick segments show interquartile ranges, dark ticks mark medians, and annotations report means.
Descriptor
Coverage
DirAcc ↑
RankCorr ↑
AlignedGap ↑
(%)
(%)
(pp)
Proprio
3.6
88.2
0.650
+23.4
[75.7,97.4]
[13.6,33.0]
Vision
4.2
87.2
0.595
+29.0
[71.4,100.0]
[19.6,38.5]
Fusion
3.8
93.9
0.758
+33.7
Appendix
Table 9: Independent descriptor comparison. Coverage is the percentage of the 20,480 corpus chunks receiving nonzero credit; quality metrics use descriptor-specific MC samples and the definitions in Appendix D.3 . Brackets give 95% rollout-group-cluster bootstrap intervals.
Credit weight λ
0.1
0.2
0.3
0.4
Success (%)
95.8
97.0
96.5
96.9
Appendix
Table 10: Credit-weight sensitivity on LIBERO-Spatial. Task success is evaluated after 63 training steps; bold marks the default weight.
δedge
Retained
DirAcc (%) ↑
AlignedGap (pp) ↑
0.05
25
100.0
+30.8
0.10
30
100.0
+29.3
0.15
45
97.5
+28.5
0.30
46
95.1
+27.3
Appendix
Table 11: Gate-parameter sensitivity on 160 locked chunks. Retained counts nonzero credits; bold marks the default setting. Metric definitions are given in Appendix D.3 .
η
Credited chunks
Nodes
Summaries
Evictions
0.90
628
287
873
0
0.93
777
484
1,405
0
0.96
604
1,234
3,064
0
0.99
302
7,458
12,325
4,678
Appendix
Table 12: Credit coverage and archive size under frozen-policy replay of the same 20,480 chunks. Credited chunks counts nonzero gated credits; Nodes and Summaries are final archive totals across tasks. Evictions counts cumulative node removals due to the 1,024-node limit per task.
Method
LIBERO-10 success (%)
Fast-WAM (official full SFT)
95.2
SFT initialization (step 6,000)
53.4
+ GRPO
92.8
+ PRICE
97.2
Appendix
Table 13: Success rates (%) on LIBERO-10 with Fast-WAM, evaluated over 50 episodes per task. The official full-SFT result is included for reference.
Figure 7: Additional simulated execution examples on LIBERO-Spatial. Each row shows five chronological frames from one recorded rollout: successful bowl transfers from (a) the tabletop, (b) a drawer, and (c) the stovetop, followed by (d) an unsuccessful tabletop transfer. Labels f are zero-based video frame indices.
Figure 8: Additional real-robot execution examples. (a) Remote Insertion and (b) Produce Sorting each show one successful episode (upper row) and one unsuccessful episode (lower row), with five chronological frames per episode. Timestamps are relative to episode start. Remote Insertion uses the lower left-wrist camera; Produce Sorting uses the lower right-wrist camera to show both trays. Outcomes are operator-provided episode labels.
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verifier outcome as a uniform advantage over all action tokens. This outcome signal is useful but structurally incomplete: it punishes useful exploration in failed rollouts and reinforces redundant or regressive actions in successful rollouts. We propose TRIAGE, a role-typed credit assignment framework that adds a semantic role axis to outcome credit. A structured judge classifies each segment as decisive progress, useful exploration, no-progress infrastructure, or regression, and a fixed role-conditioned rule maps these labels to bounded segment-level process rewards. This keeps verifier outcomes as the source of optimization direction while correcting the two main blind spots of outcome-only credit. We further show that role-conditioned credit is the optimal segment-level correction expressible from role labels alone -- a projection of the per-segment advantage residual onto the role variable -- so that the fixed role constants reduce advantage estimation error whenever the judge is reliable, and we connect this to lower-variance policy gradients. Across ALFWorld, Search-QA, and WebShop, TRIAGE improves success rates over GRPO for two policy models and outperforms both a scalar judge-derived process reward and an outcome-supervised shared-backbone value baseline. Ablations show that the gain comes from role typing rather than merely adding dense rewards: reliable detection of regression inside successful trajectories is the dominant contributor, while exploration credit provides a consistent secondary gain; on completed ALFWorld and WebShop rollouts, TRIAGE also reduces environment-facing turns by an additional 10.4% and 14.8% relative to GRPO.
Yuanda Xu, Zhengze Zhou, Hejian Sang +6
1LinkedIn Corporation · 2Harvard University · 3Johns Hopkins University +1
Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields no training signal, failed attempts cannot be told apart by how close they came to completion, and turns that advance the task receive the same credit as turns that only query the environment. Prior work refines the unit of comparison from the trajectory to the step, or trains a reward model to supply intermediate signal: the former still derives its signal from final success alone, and the latter estimates it with a model. We observe that the acceptance checks that decide success can also be run on intermediate states, so progress is as verifiable as the outcome. We propose ProCredit, which turns this verified progress into credit: it reruns the acceptance checks after each turn, rewards the turn by its change in progress, and uses these rewards to assign credit both across attempts at the same task and across the turns within a trajectory. Starting from Qwen3.5 base models at three scales on AppWorld, ProCredit outperforms outcome-reward baselines and progress-based baselines in task completion rate at every scale on both test sets, exceeding the strongest outcome-reward baseline by 4.1 percentage points at 4B, and results in a second environment show the same direction of improvement. Ablations show that adding the final progress to the trajectory score alone does not improve performance: the gain comes from crediting progress to the turn where it occurs.
Ming Ma, Yi Zhu, Yiran Zhong +7
Institute of Neuroscience, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Tongyi Lab, Alibaba Group +2
Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation offers finer supervision by re-evaluating sampled behavior with privileged information (PI) available only during training. However, fine-grained supervision is not necessarily fine-grained credit: PI-induced likelihood changes describe how additional information alters policy preference, but do not directly determine how an executable action should inherit the verified task outcome. This creates a supervision-credit gap. Privileged signals may be irrelevant to the current interaction state, operate at a token granularity misaligned with executable decisions, and lack the outcome semantics required for reinforcement. We introduce TASPO, which converts privileged supervision into outcome-grounded action credit. TASPO constructs decision-applicable PI from verified successful experience, aggregates PI-induced likelihood shifts at the executable-action level, and converts relative action support into positive, bounded, mean-preserving weights on the original trajectory advantage. Thus, the verified outcome determines the update direction and average scale, while PI only redistributes credit across actions. Across three agentic benchmarks, TASPO improves over GRPO by 10.6% and generalizes better to unseen tasks. Further analysis indicates that TASPO reduces supervision mismatch and that action-level assignment stabilizes the policy optimization process. These findings offer the community another interesting perspective.