Disentangling Spurious Correlations in Vision-Language-Action Models via Predicting Domain-Invariant Latent Lookahead
Authors: Junghyun Kim, Ngseo Kim, ChungWoo Lee, Seoyeon Lee, Woo-Jeong Baek, Adam Zhou, Chip Huyen, Jun-Ki Lee, +2 more
Organizations: OpenMind, San Francisco, CA, USA · Seoul National University, Seoul, Korea · Hyundai Motors, Korea · Ajou University, Korea · Tommoro Robotics, Korea
Vision-Language-Action (VLA) models remain brittle under visual distribution shifts, often relying on spurious correlations tied to domain-specific factors rather than task-relevant structure. We propose Domain-Invariant Latent Lookahead (DILL), a representation-learning framework that mitigates shortcut learning in VLA policies. Our key idea is to supervise policies with domain-invariant future latents learned from domain-transformed trajectory data. A Task-Domain Encoder is trained with contrastive objectives and Gaussian disentanglement regularization to separate task-relevant structure from domain-specific visual variation. The learned encoder then provides future latents for VLA policy learning through lookahead prediction and domain disentanglement, encouraging the policy to focus on task-relevant structure rather than incidental visual factors. Counterfactual task-view evaluations show that DILL reduces shortcut reliance, while LIBERO-Plus evaluations demonstrate improved visual robustness, with 69.1% average success, 11.4 percentage points above the strongest baseline. Real-world manipulation experiments further support DILL's applicability beyond controlled simulation. Complementary latent-space diagnostics show that these behavioral gains are accompanied by representations that better preserve task-consistent structure while suppressing domain-specific variation. Our project page is available at https://dill-vla.github.io/.
Figures & tables
Figure 1: Shortcut learning under task–domain confounding. When tasks and visual domains are spuriously correlated, a VLA may use domain-specific appearance as a shortcut for action selection. DILL reduces this failure by conditioning actions on a domain-invariant latent lookahead.
Figure 2: Overview of Domain-Invariant Latent Lookahead. We first learn a Task-Domain Encoder that maps observation chunks into task and domain latents. During policy learning, a Lookahead Predictor predicts a future task latent from the current observation and instruction, while a Current Representation Head produces a domain-disentangled current representation. The two policy latents condition the action head for control.
Figure 3: LIBERO shortcut diagnostic.
Model
Original
Visual perturbations
Average
Camera
Light
BG
Noise
OpenVLA [ 21 ]
76.5
0.8
8.1
34.8
15.2
14.7
↓ 75.7
↓ 68.4
↓ 41.7
↓ 61.3
↓ 61.8
WorldVLA [ 9 ]
79.1
0.1
43.7
17.1
10.9
18.0
↓ 79.0
↓ 35.4
↓ 62.0
↓ 68.2
↓ 61.2
UniVLA [ 8 ]
95.5
1.8
69.0
81.0
21.2
43.3
Table 1: Zero-shot robustness evaluation on visual perturbations in LIBERO-Plus [ 13 ] . For each model, the first row reports success (%) and the second its drop from Original in percentage points. Average is the unweighted mean across the four visual categories. The bottom block matches policy architecture and augmented source data (SA: source augmentation).
Figure 4: Pairwise similarity diagnostics for learned latents. We compare cosine-similarity distributions for task pairs and domain pairs constructed from unseen task–domain combinations. From left to right, the panels show the task encoder, domain encoder, and final VLA latent. The task encoder should group task pairs, the domain encoder should group domain pairs, and the final VLA latent should preserve task-consistent structure while suppressing domain-specific variation.
Shortcut diagnosis
Robustness
Model
CTF ↑
SD ↓
No pert. ↑
Predefined ↑
New ↑
Base VLA
29.7±5.4
47.2±4.8
84.1±5.5
48.3±5.9
64.0±6.4
DILL-Current
83.3±4.8
0.0±0.0
88.9±2.7
78.2±7.6
77.2±5.1
DILL
80.2±5.5
0.0±0.0
88.9±2.7
83.8±4.2
82.0±4.0
Table 2: Real-world shortcut diagnosis and robustness (%). CTF measures command-following under swapped target-color/viewpoint pairings; SD is shortcut degree (Section 4.1 ). Predefined perturbations are held-out instances of encoder-training transformation families; new perturbations (cast shadows, dynamic backgrounds, and foreground clutter) are absent from pair construction.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Frames
Trajectories
Tasks
MimicGen, 16 datasets
1,010,618
3,200
16
ManiSkill, 2 datasets
1,632,291
14,057
17
OXE, Bridge + Fractal/RT-1
5,785,810
140,404
20,553
Total
8,428,719
157,661
20,586
Appendix
Table 3: Source trajectory collection used for Task-Domain Encoder pretraining. The MimicGen row aggregates 16 task datasets, while OXE aggregates Bridge and Fractal/RT-1 subsets.
Family
Transformations
Shared condition for domain positives
Viewpoint changes
Rendered camera view, random crop, intrinsics change, radial distortion, warping
Camera ID or camera parameters; crop seed; focal scale; distortion and warp parameters
Environment appearance
Background replacement, lighting changes, color changes
Background texture IDs; mask regions; replacement mode; lighting or color parameters
Camera heterogeneity
Camera unprocessing, sensor noise, camera color response changes
Camera-pipeline seed; color matrices; gamma; ISO level; shot/read noise; channel gains
Table 4: Domain transformation families used for positive pair mining. A domain condition η consists of a transformation family and its sampled parameters or random seed. The last column lists the condition shared by domain-positive pairs.
Figure 5: Representative single-family task-positive transformations. Each column after the anchor shows the same trajectory chunk under one family of domain-specific visual factors.
Figure 6: Representative assets for mask-based background replacement. We sample table and floor textures from Poly Haven and generate wall and background assets with Stable Diffusion XL.
Module
Parameters
Lookahead Predictor, AttentiveLatentHead
167.85M
Current Representation Head, AttentiveLatentHead
167.85M
ResNetActionHead, dual-head input
28.48M
Total policy-side heads
364.18M
Appendix
Table 5: Policy-side head budget for the VLA implementation.
Stage
Data
Updated modules
Objective
Task-domain encoder pretraining
Source trajectory collection
Task/domain encoder LoRA adapters and projection heads
LTDD
VLA policy pretraining
Source trajectory collection
VLM LoRA, Lookahead Predictor, Current Representation Head, action head
VLM LoRA, Lookahead Predictor, Current Representation Head, action head
Lpolicy
Appendix
Table 6: Training Stage Summary for Domain-Invariant Latent Lookahead.
Metric
Base VLA
DILL
Downstream tuning cost ( × Base VLA)
1.00
1.42
Peak training memory (GB)
20.9
22.4
Inference latency (ms per 30-action chunk)
69.9
70.2
Appendix
Table 7: Training and inference costs. The one-time Task-Domain Encoder pretraining cost is reported separately in the text. Latency is measured in a separate benchmark with 30-action chunks.
Variant
Components
Metrics
SA
TD
CH
LA
In-dist. SR ↑
Center OOD SR ↑
Counter OOD SR ↑
Avg. SR ↑
Shortcut degree ↓
Base VLA
68%
22%
0%
30%
73%
Base VLA + SA
√
82%
10%
2%
31%
69%
Entangled LA
√
√
76%
24%
0%
33%
65%
DILL w/o CH
√
√
√
60%
44%
44%
49%
6%
DILL w/o LA
√
√
√
50%
62%
62%
58%
3%
Appendix
Table 8: Ablation study. All variants are evaluated on the LIBERO shortcut diagnostic under the same task–view island protocol. Component columns indicate whether each variant uses source augmentation (SA), task–domain latent targets from the task/domain encoders (TD), the disentangled current head (CH), and latent lookahead alignment (LA). In-dist. SR averages the original training compositions, Center OOD SR evaluates unseen midpoint viewpoints without swapping task identity, and Counter OOD SR evaluates counterfactual task–view swaps. Shortcut degree measures how often the policy follows the task associated with the observed view rather than the commanded task. Blank component entries indicate that the component is absent.
Model
Original
LIBERO-Plus perturbations
Total
Camera
Robot
Language
Light
BG
Noise
Layout
OpenVLA [ 21 ]
76.5
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
↓ 75.7
↓ 73.0
↓ 53.5
↓ 68.4
↓ 41.7
↓ 61.3
↓ 48.0
↓ 60.9
WorldVLA [ 9 ]
79.1
0.1
27.9
41.6
43.7
17.1
10.9
38.0
25.0
↓ 79.0
↓ 51.2
↓ 37.5
↓ 35.4
↓ 62.0
↓ 68.2
↓ 41.1
↓ 54.1
UniVLA [ 8 ]
95.5
1.8
46.2
69.6
69.0
81.0
21.2
31.9
43.9
Appendix
Table 9: Full LIBERO-Plus robustness breakdown. For each model, the first row shows success rate (%) on the original LIBERO benchmark and all seven LIBERO-Plus perturbation categories. The second row shows the absolute drop relative to the original score. Total denotes the official LIBERO-Plus leaderboard score.
Figure 7: Pair construction for latent similarity diagnostics. Task pairs test whether a latent remains stable across visual-domain changes when the underlying trajectory content is preserved. Domain pairs test whether a latent collapses observations that share visual-domain cues despite different trajectory content.
Figure 8: Final action-conditioning latent similarity diagnostics. We compare Base VLA and DILL using the final latent representation provided to the action head. A task-centric action-conditioning latent should assign higher similarity to task pairs than to domain pairs.
Representation
Task-pair similarity ↑
Domain-pair similarity ↓
Gap Δtask↑
Base VLA final latent
0.52
0.70
−0.18
DILL final latent
0.94
0.55
0.39
Appendix
Table 10: Summary of final action-conditioning latent similarity. The gap Δtask is computed as mean task-pair similarity minus mean domain-pair similarity. Larger gap indicates a more task-centric action-conditioning representation.
Figure 9: Real-world counterfactual color–viewpoint compositions. Training pairs red-target instructions only with the left view and blue-target instructions only with the right view, creating a spurious correlation between target color and viewpoint. Counterfactual testing swaps these pairings while preserving the instruction and desired manipulation. Each row shows five frames from a training demonstration and one example of the held-out test view.
Figure 10: Real-world visual robustness tasks and evaluation conditions. (a) Five ordered frames from the training demonstration for each task: shoe upright placement, tissue pulling, and laptop closing. (b) Representative evaluation scenes: no perturbation; viewpoint and background changes from predefined transformation families; and dynamic backgrounds and foreground clutter from new families absent from Task-Domain Encoder pair construction. All three tasks are evaluated under all three conditions, with their instructions and manipulation goals unchanged. Quantitative results are reported in Table 2 .
Vision-language-action (VLA) models that generate continuous action chunks via flow matching lack an internal signal for judging whether a given prediction is reliable. Distribution shift and long-horizon rollouts can push backbone representations away from the region the action head decodes reliably, yet the policy has no mechanism to detect or react to this drift. We observe that the cost of transporting observation features to the action representation in a shared feature space rises precisely when such drift occurs, providing a per-step reliability estimate without extra supervision. Building on this observation, we propose DiG (Discrepancy Gate), a lightweight plug-in module for flow-matching VLA policies. DiG computes a sliced Wasserstein transport cost between backbone features and the action expert's own input projection, maps it through an exponential gate, and uses the gate to modulate both a residual feature refinement and the training loss. At inference time, the gate enables DiG-Refinefine, an iterative refinement process that corrects action chunks before execution. Experiments on both simulation and real-world scenarios show that DiG consistently improves success rates, with the largest gains under distribution shift and on long-horizon tasks.
Wanpeng Zhang, Ye Wang, Hao Luo +6
1Peking University · 3BeingBeyond · 2Renmin University of China
Future prediction is increasingly used to improve vision-language-action (VLA) policies, based on the premise that anticipating scene evolution encourages representations useful for control. However, forecast quality alone does not establish that a policy has learned a better representation for action. This distinction matters under distribution shift, where successful control depends on preserving spatial state and likely scene change beyond familiar configurations. We study what determines whether predictive supervision improves the visual representation used by a VLA policy. Through controlled comparisons with matched target constructions, prediction horizons, and training conditions, we find that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond familiar scenes. We trace these differences to how predictive errors shape the policy's visual stream. Consistent with this controlled finding, VLA policies trained with more direct, scene-matched future supervision show stronger robustness under simulated and physical distribution shifts. Together, our results frame future prediction as a representation-learning design problem whose value for control depends on whether its supervision reaches the representations through which the policy acts.
Hanseul Kim, Jewon Yeom, Youngjoon Jeong +2
Graduate School of Data Science Seoul National University
Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varying environments. While existing approaches attempt to construct behavior representations through action-centric latent variables, they are often limited by short-horizon temporal fragmentation and static execution-alignment, leading to inconsistent behaviors in complex scenarios. To address these limitations, we propose \textbf{BehaviorVLA}, a framework that facilitates robust manipulation through the learning of a temporally coherent behavioral representations. Our approach features two symmetric components: (1) the \textbf{Visuomotor Behavior Encoder (VBE)}, which utilizes a causal Mamba-based architecture to aggregate long-horizon trajectory information into a unified behavior representation; and (2) the \textbf{Phase-conditioned Behavior Decoder (PBD)}, which decodes this representation into precise actions by dynamically aligning task-level priors with real-time execution progress. Experiments on RoboTwin 2.0, LIBERO, and CALVIN demonstrate state-of-the-art success rates of 58%, 98%, and 4.36 (Avg.Len), respectively. Notably, in real-world sim-to-real transfer, BehaviorVLA matches the performance of OpenVLA-OFT using only 50% of the demonstration data, showcasing its superior data efficiency and generalization.
Bing Hu, Zaijing Li, Rui Shao +4
Harbin Institute of Technology, Shenzhen · PengCheng Laboratory · Shenzhen Loop Area Institute +2