Organizations: School of Data Science, Fudan University · Shanghai Innovation Institute · Beihang University · Simple AI · Tsinghua University · Tuojing Intelligence · The University of Hong Kong
Outcome-based reinforcement learning (RL) post-trains vision--language--action policies using terminal success signals, but assigns the same trajectory-level advantage to every action chunk. A failed episode can thus penalize useful early actions as if they caused the failure. Existing approaches seek finer-grained feedback through learned evaluators, adding task-specific supervision or additional model training. We explore, for the first time to our knowledge, whether physical relations across trajectories can provide action-chunk credit in embodied RL from terminal outcomes alone, without an auxiliary evaluator. The key insight is that rollouts reaching corresponding physical situations can serve as references for one another: their terminal outcomes provide evidence for assessing local progress. We introduce Physical Relations for Inferring Credit from Episodes(PRICE), with two components: (i) a physical relational graph that pools current and historical outcomes at corresponding chunk boundaries to estimate success potentials; and (ii) confidence-gated credit assignment that uses changes in these potentials to refine trajectory-level supervision. Our analysis connects oracle potential changes to the terminal-success objective and provides a finite-sample directional bound for outcome-independent evidence pools. Independent continuation tests show that PRICE's retained credits align with local progress, while experiments on LIBERO, RoboTwin 2.0, and real robots demonstrate improved task success over outcome-based baselines and faster learning.
Figures & tables
Figure 1: Motivation for PRICE. PRICE derives local action-chunk credit by comparing physically corresponding states across rollouts, using only terminal outcomes and no auxiliary evaluator.
Figure 2: Overview of PRICE. (a) Visual–proprioceptive descriptors match physically corresponding states across rollouts. (b) Cross-trajectory terminal outcomes determine node success potentials. (c) Confidence-gated endpoint differences provide action-chunk credit that augments the GRPO advantage for policy optimization.
Method
Policy input
Spatial
Object
Goal
Long
Avg.
VLA baseline
OpenVLA ( Kim et al., 2024 )
T+I
84.7
88.4
79.2
53.7
76.5
OpenVLA ∗ -Full ( Fei et al., 2026 )
T+I
91.6
95.3
90.6
86.5
91.0
Evaluator-based RL
TGRPO ( Chen et al., 2025b )
T+I
90.4
92.2
81.0
59.2
80.7
GRAPE ( Zhang et al., 2024 )
T+I
88.5
92.1
83.1
57.2
80.2
Table 1: Performance comparison on the LIBERO benchmark. We report task success rates (%). Rows below the dashed line are our results; other values follow the unified comparison of Fei et al. (2026) . T, W, P, and I denote third-person images, wrist images, proprioception, and language instructions, respectively. Full and One denote full-shot and one-trajectory-per-task SFT, respectively.
Method
Handover Block
Lift Pot
Move Can Pot
Pick Dual Bottles
Place Container Plate
Place Empty Cup
Avg.
SimpleVLA-RL ( Li et al., 2026a )
57.8
64.1
61.2
68.3
82.1
94.2
71.3
RLinf-VLA ( Zang et al., 2025 )
70.31
70.31
83.59
92.96
95.31
94.53
84.5
Feat2Go ( Shu et al., 2026 )
76.6
78.9
91.4
93.8
96.9
95.3
88.8
OpenVLA-OFT (SFT)
28.1
3.1
9.4
20.3
54.7
75.8
31.9
+ PRICE (Ours)
79.7
80.5
90.6
94.5
97.7
96.9
90.0
↑ 51.6
↑ 77.4
↑ 81.2
↑ 74.2
↑ 43.0
↑ 21.1
↑ 58.1
Table 2: Performance comparison on six selected RoboTwin 2.0 manipulation tasks. We report task success rates (%) and the average across the six tasks. The OpenVLA-OFT (SFT) row is the initialization used for PRICE; results for the other compared methods are from the cited works.
Figure 3: Training efficiency on LIBERO. (a–b) Rollout success during training on Goal and Spatial. Solid curves show seven-step moving averages and faint lines show raw observations; dashed guides and x-axis ticks mark the steps for 90% and 95% success. (c) Training steps of SRPO and PRICE across four suites. SRPO counts are from Fei et al. (2026) ; corresponding success rates are reported in Table 1 .
Figure 4: Real-robot performance and representative executions. (a) Success rates of the initialization, GRPO-AWR, and PRICE-AWR on Remote Insertion and Produce Sorting under the HiFi-UMI evaluation protocol. (b) Chronological keyframes show remote transport, alignment, and insertion into the storage box (top), and fruit grasping, transport, and release into its designated tray during Produce Sorting (bottom).
Variant
History
Gate
Spatial
Object
Avg.
GRPO
×
×
93.2
94.6
93.9
PRICE w/o Historical Evidence
×
✓
93.4
94.6
94.0
PRICE w/o Gate
✓
×
95.8
97.9
96.9
Full PRICE
✓
✓
97.0
99.6
98.3
Table 3: Component ablations on LIBERO-Spatial and LIBERO-Object. All variants use 63 training steps with the same initialization, interaction budget, and optimization settings within each suite. We report task success (%) and the average across the two suites.
Figure 5: Credit analysis on LIBERO-Spatial. Rows denote trajectory windows and columns action chunks. (a) GRPO advantage sign. (b) Gated PRICE credit from (e). (c) Independent MC progress over the pooled evidence in (e). (d–e) Pre-gate contrasts without/with history at fixed nodes. (f) MC-based credit evaluation for all nonzero candidates in (e) and their rejected/retained subsets.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Setting
Parameter
Setting
Simulation training and GRPO
PRICE in simulation
Optimizer
AdamW
Credit weight λ
0.2
Learning rate
5×10−6
Matching threshold η
0.93
Learning-rate schedule
Constant
Gate parameter δedge
0.15
Cases per raw rollout batch
64
Summaries per node Chist
4
Rollouts per case
8
Nodes per task Cnode
1,024
Appendix
Table 4: Core training hyperparameters. Optimization batch sizes count trajectories and are global across eight GPUs. Simulation and real-robot credit weights are reported separately.
LIBERO
Method
Spatial
Object
Goal
Long
Avg.
PRICE
99.2±0.2
99.6±0.2
99.4±0.3
97.6±0.5
99.0±0.2
Appendix
Table 5: Main PRICE results (%, mean ± SD) across three independent training runs.
Suite
Success (%)
SimpleVLA-RL
PRICE
Reduction (%)
Goal
90
441.9
272.8
38.3
95
625.2
517.2
17.3
Spatial
90
329.0
221.2
32.8
95
714.6
544.3
23.8
Appendix
Table 6: Cumulative GPU hours at the marked milestones in Figure 3 (a–b). Reductions are relative to SimpleVLA-RL.
Training round
DirAcc (%) ↑
RankCorr ↑
AlignedGap (pp) ↑
16
95.8
0.67
+24.6
32
95.8
0.71
+27.8
63
94.3
0.68
+26.2
Appendix
Table 7: Online credit alignment on LIBERO-Spatial. Of 1,200 audited chunks per stage across three PRICE runs, 24, 48, and 36 receive nonzero credit at rounds 16, 32, and 63, respectively.
Selector
DirAcc (%) ↑
AlignedGap (pp) ↑
Confidence gate
97.5
+28.5
Endpoint-count selection
65.9
+4.6
Random selection
77.2
+12.8
Appendix
Table 8: Credit selection at equal coverage on LIBERO-Spatial. Each selector retains 45 of the same 123 candidates. MC ties are excluded from DirAcc; the random row reports means over 100,000 subsets.
Figure 6: Credit selection and local progress on LIBERO-Spatial. (a) Raw PRICE credit versus MC progress for the same 123 candidates as Figure 5 ; blue and gray denote the 45 retained and 78 rejected candidates, respectively. Shaded quadrants indicate agreement in sign. (b) Per-chunk aligned MC gaps ge (pp), whose means give AlignedGap. Vertical jitter separates points; thick segments show interquartile ranges, dark ticks mark medians, and annotations report means.
Descriptor
Coverage
DirAcc ↑
RankCorr ↑
AlignedGap ↑
(%)
(%)
(pp)
Proprio
3.6
88.2
0.650
+23.4
[75.7,97.4]
[13.6,33.0]
Vision
4.2
87.2
0.595
+29.0
[71.4,100.0]
[19.6,38.5]
Fusion
3.8
93.9
0.758
+33.7
Appendix
Table 9: Independent descriptor comparison. Coverage is the percentage of the 20,480 corpus chunks receiving nonzero credit; quality metrics use descriptor-specific MC samples and the definitions in Appendix D.3 . Brackets give 95% rollout-group-cluster bootstrap intervals.
Credit weight λ
0.1
0.2
0.3
0.4
Success (%)
95.8
97.0
96.5
96.9
Appendix
Table 10: Credit-weight sensitivity on LIBERO-Spatial. Task success is evaluated after 63 training steps; bold marks the default weight.
δedge
Retained
DirAcc (%) ↑
AlignedGap (pp) ↑
0.05
25
100.0
+30.8
0.10
30
100.0
+29.3
0.15
45
97.5
+28.5
0.30
46
95.1
+27.3
Appendix
Table 11: Gate-parameter sensitivity on 160 locked chunks. Retained counts nonzero credits; bold marks the default setting. Metric definitions are given in Appendix D.3 .
η
Credited chunks
Nodes
Summaries
Evictions
0.90
628
287
873
0
0.93
777
484
1,405
0
0.96
604
1,234
3,064
0
0.99
302
7,458
12,325
4,678
Appendix
Table 12: Credit coverage and archive size under frozen-policy replay of the same 20,480 chunks. Credited chunks counts nonzero gated credits; Nodes and Summaries are final archive totals across tasks. Evictions counts cumulative node removals due to the 1,024-node limit per task.
Method
LIBERO-10 success (%)
Fast-WAM (official full SFT)
95.2
SFT initialization (step 6,000)
53.4
+ GRPO
92.8
+ PRICE
97.2
Appendix
Table 13: Success rates (%) on LIBERO-10 with Fast-WAM, evaluated over 50 episodes per task. The official full-SFT result is included for reference.
Figure 7: Additional simulated execution examples on LIBERO-Spatial. Each row shows five chronological frames from one recorded rollout: successful bowl transfers from (a) the tabletop, (b) a drawer, and (c) the stovetop, followed by (d) an unsuccessful tabletop transfer. Labels f are zero-based video frame indices.
Figure 8: Additional real-robot execution examples. (a) Remote Insertion and (b) Produce Sorting each show one successful episode (upper row) and one unsuccessful episode (lower row), with five chronological frames per episode. Timestamps are relative to episode start. Remote Insertion uses the lower left-wrist camera; Produce Sorting uses the lower right-wrist camera to show both trays. Outcomes are operator-provided episode labels.