WorldGuide: Learning Success-Failure Boundaries in Latent World Models for Vision-Language-Action Policies
Authors: Lin Liu, Lu Zhang, Ziying Song, Wu Yang, Yuzheng Zhuang, Yunzhi Zhuge, Shuai Tao, Wulong Liu, +1 more
Organizations: School of Information and Communication Engineering Dalian University of Technology & Beta Infinity · School of Information and Communication, Engineering Dalian University of Technology · Nanyang Technological University · Beta Infinity
Latent world models offer a promising way to improve Vision-Language-Action policies by capturing the consequences of actions. However, models trained primarily on expert demonstrations have limited exposure to failure outcomes and may struggle to distinguish visually similar successful and failed interactions. We propose \textbf{WorldGuide}, a framework that learns these distinctions in latent space and uses them to guide policy training. WorldGuide combines predictive pretraining on successful and failed trajectories with contrastive learning on matched success--failure pairs. The learned predictor then provides a differentiable reward to guide joint optimization of the policy and visual encoder. The predictor is discarded after training, so deployment requires no additional world-model inference. Extensive experiments show that WorldGuide substantially improves VLA reliability and achieves state of the art performance on LIBERO 100 and SimplerEnv, reaching \textbf{96.8%} and \textbf{72.0%}, respectively. Code will be publicly available.
Figures & tables
Figure 1: Schematic motivation for WorldGuide. (a) Success-only training poorly constrains latent dynamics for failures. (b) WorldGuide combines failure-rich predictive learning with matched comparisons to yield outcome-sensitive representations for policy guidance.
Figure 2: Overview of WorldGuide. (a) Stage 1: Failure-Rich Pretraining learns latent dynamics from successful and failed trajectories. (b) Stage 2: Boundary-Aware Contrastive Fine-Tuning distinguishes visually matched success–failure pairs while retaining the predictive objective. (c) Stage 3: End-to-End Optimization uses a differentiable reward from the frozen world model to train the VLA policy. Flame and snowflake icons indicate trainable and frozen modules, respectively.
Figure 3: Success–failure matching for contrastive learning. Each successful segment is paired with the visually closest failure candidate from the same task within a temporal window to provide contrastive supervision; a dashed line marks the failure onset tf .
(a) LIBERO Benchmark
Method
Category
Model Size
Latency (ms)
Speedup
100
Goal
Object
Spatial
Average
OpenVLA-OFT Kim et al. (2024)
VLA
7B
277
–
94.5
97.9
98.4
97.6
97.1
π0.5 Intelligence et al. (2025)
VLA
3.5B
220
–
92.4
98.0
98.2
98.8
96.9
GR00T-N1.6 Bjorck et al. (2025)
VLA
3.3B
259
–
94.4
97.5
98.5
97.7
97.0
UniVLA Bu et al. (2025)
Latent-VLA
7B
–
–
92.0
95.6
96.8
96.5
95.2
Mantis Yang et al. (2026)
Latent-VLA
5.8B
–
–
94.2
94.4
99.2
98.8
96.7
Table 1: Quantitative Evaluation on LIBERO and SimplerEnv Benchmarks. Results are reported as percentages. Best and second best overall results are highlighted in dark green and light green, respectively.
Task
Method
Press Stapler
Move Playingcard Away
Place Object Stand
Place Container Plate
Turn Switch
Lift Pot
π0.5
0.97 / 0.95
0.94 / 0.98
0.87 / 0.87
0.95 / 0.93
0.62 / 0.69
0.99 / 0.99
π0 -Fast
0.96 / 0.97
0.95 / 0.98
0.86 / 0.92
0.92 / 0.98
0.63 / 0.67
1.00 / 0.99
GR00T-N1.5
0.98 / 0.98
0.99 / 0.99
0.97 / 0.96
0.85 / 0.89
0.58 / 0.64
0.99 / 1.00
StarVLA-OFT
0.97 / 0.96
1.00 / 0.98
0.84 / 0.85
0.90 / 0.97
0.67 / 0.60
1.00 / 1.00
WorldGuide
1.00 / 0.97
0.99 / 1.00
0.94 / 0.93
0.99 / 0.97
0.66 / 0.69
0.99 / 1.00
Table 2: Comparison on RoboTwin 2.0 across 12 tasks. Each entry reports the success rate under the clean and the domain-randomized setups as x/y .
Purple square
Green circle
Yellow triangle
Avg.
π0
10%
20%
5%
11.7%
π0.5
10%
15%
10%
11.7%
WorldGuide
20%
20%
15%
18.3%
Table 3: Real-world results on the ARX LIFT2. Success rate (%) over 20 trials per task per method
Readout / Evaluation Setting
Stage 1
Stage 2
Window AUC (decision-critical)
0.74
0.79
Episode AUC
0.81
0.87
Phase: Early ( p<0.35 )
0.63
0.60
Phase: Mid ( 0.35≤p<0.70 )
0.74
0.82
Phase: Near-failure ( p≥0.70 )
0.75
0.85
Phase: Post-failure ( τ≥0 )
0.73
0.77
Table 4: Feature separability under Stage 2 boundary-aware fine-tuning. Episode-held-out linear probing AUCs (Fig. 4 ) before (Stage 1) and after (Stage 2) fine-tuning. Middle rows detail AUC across task progress p and failure offset τ . Bottom rows evaluate early-warning generalization: training the probe strictly on near-failure frames ( τ∈[−150,0) frames) and testing transferability to unseen future frames.
Figure 4: Failure success separability of the world model predictor.
Setting
Stage
RoboTwin
LIBERO
SimplerEnv
1
2
Clean
Randomized
Google
WidowX
Base VLA
✗
✗
80.1%
80.2%
96.5%
68.8%
60.2%
Stage 1 (success only)
✓
✗
82.4%
81.9%
97.5%
69.6%
63.0%
Stage 1
✓
✗
83.0%
83.5%
97.8%
70.7%
63.4%
Stage 1 + 2
✓
✓
85.7%
86.1%
98.4%
72.0%
64.8%
Table 5: Stage-wise ablation.
Figure 5: Reward-source ablation .
Figure 6: Fine-grained manipulation comparison between π0.5 and WorldGuide in real world.
Figure 7: Attention weight matrix of latent action tokens to image tokens.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
LIBERO
RoboTwin 2.0
SimplerEnv
VLM backbone
Qwen3-VL-4B
Qwen3-VL-4B
Qwen3-VL-4B
Action head
OFT tokens
OFT tokens
DiT-B flow matching
Action space
Δ qpos, 7-DoF
abs qpos, 14-DoF
Δ EE, 7-DoF
Action chunk H
8
50
16
WM encoder ϕθ
ViT-L/14, 2242
ViT-L/14, 2242
ViT-T/14, 2242
WM window (ctx + pred)
8+1 , frameskip 1
50+1 , frameskip 1
16+1 , frameskip 1
Appendix
Table 6: Architecture and per-benchmark configurations. The world model is discarded at deployment; only the VLA policy executes.
Stage 1 (pretrain)
Stage 2 (contrastive)
Stage 3 (reward)
LIBERO
Steps
60,000
10,000
50,000
Batch / device
16
16
16
Learning rate
5×10−5 (WM)
5×10−5 (WM)
1×10−5 (VLA)
LR schedule
cosine →1×10−6
cosine →1×10−6
cosine →2×10−6
Warmup ratio
0.1
0.1
0.1
Appendix
Table 7: Stage-wise training configurations across different environments.
Figure 8: Real world experimental setup: ARX Lift 2 with PICO.
Figure 9: Success–failure separability on RoboTwin stack_blocks_three . (a) Single-window failure-score distributions. (b) Episode mean scores. (c) Stage-wise AUC. (d) Boundary score swept along the time offset τ from the expected task end.
Variant
RoboTwin
LIBERO
Counterpart selection (separated with dm )
Random opposite-outcome
78.3/77.1
95.5
Progress-anchored, random pick
84.7/84.2
98.0
Separation metric (matched counterparts)
Uniform α
85.0/84.6
98.1
g≡1
85.3/85.0
98.2
Appendix
Table 8: Stage-2 design choices. Policy success rates (%) on RoboTwin (clean/randomized placements) and LIBERO with one Stage-2 component replaced at a time; metric variants replace dm in both the contrastive loss and the Stage-3 reward.
Figure 10: Hyper-parameter sensitivity. Success rate on RoboTwin (50 episodes per task, clean and randomized splits) as a function of reward weight β , contrastive weight λc , hinge margin m , match tolerance Δ , and prediction horizon W , with all other hyper-parameters fixed at defaults (ringed marker, dashed line). The two loss weights show opposite asymmetries: under-weighting the boundary reward ( β ) is nearly harmless, whereas over-weighting it reduces performance by 3.3 points; conversely, shrinking λc removes Stage 2 entirely. The hinge margin m acts as a threshold, saturating beyond 0.35 . Match tolerance Δ is comparatively flat, losing 2.2 points on the randomized split at Δ=6 , where narrow windows fail to retrieve counterparts and starve the contrastive loss, while excessively wide windows mix execution phases. Horizon W declines slowly from W=12 to 100 ( −0.6 points) because distant frames carry less decisive evidence; we set W=50 to align with the policy’s action chunk.
Figure 11: Learned weighting inside dm . (a) Temporal salience αt over the 50-frame context window. For near-failure windows, attention concentrates on the final ∼15 frames preceding the boundary, whereas success windows remain nearly uniform, demonstrating that the model learns when to attend. (b) Sorted channel gate gj . The gate maintains a decisive core of ∼6% of the 2048 feature channels near saturation while suppressing the rest, indicating that failure evidence resides in a low-dimensional subspace rather than being distributed across the representation.
Figure 12: Attention weight matrix of latent action tokens to image tokens on SimplerEnv.
Figure 13: Attention weight matrix of latent action tokens to image tokens on Robotwin.