Organizations: University of Science and Technology of China · Eastern Institute of Technology, Ningbo · Hangzhou Dianzi University · ShanghaiTech University
Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from 60.9% to 68.5% over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches 72.7%; on a real robot, it retains 70%--75% success under background, height, and object shifts where the base policy collapses to 0%.
Figures & tables
Figure 1: We present Juno, a unified framework that makes predictive latents useful across the vision-language-action (VLA) lifecycle. Juno integrates control-aligned JEPA pretraining, visual feature fusion and decoupled future-latent distillation, and dynamics-first test-time adaptation that updates the world model before re-aligning the policy. It improves performance on SimplerEnv and RoboCasa and supports robust real-world manipulation under distribution shifts.
Figure 2: (a) Patch-wise temporal cosine distance between frames t⋆ and t⋆+Δt for LeWM and V-JEPA2. (b) Temporal state similarity cos(ct,ct+Δt) over the episode, spanning slight motion ( t≤14 ) and pick-and-place ( t>14 ). LeWM (CLS) stays nearly constant, whereas LeWM+ Ldyn becomes sensitive during pick-and-place; mean-pooled baselines are shown for reference.
Figure 3: Overview of Juno. Top left: Dynamic CLS pretraining aligns CLS transitions with motion-weighted patch changes from a frozen JEPA supervisor. Bottom left: JEPA features enrich visual tokens, while a decoupled predictive branch distills future latent states to jointly guide action generation. Right: Test-time adaptation updates the world model using both successful and failed executions, then re-aligns the policy on verified successful executions. Flame and snowflake icons indicate trainable and frozen modules, respectively.
Method
PnP + Close
PnP From Cuttingboard
PnP From Placemat
PnP From Plate
PnP From Tray
Avg.
GR00T-N1.6
24.2
56.9
51.9
57.6
55.1
47.6
Qwen3PI
42.3
46.0
43.5
44.0
44.0
43.9
Qwen3OFT
43.7
50.4
41.5
61.0
49.2
48.8
Qwen3FAST
35.0
50.4
33.5
45.0
32.0
39.0
Qwen3GR00T
50.3
52.8
38.0
58.5
39.2
47.8
Juno (Ours)
57.3
64.8
55.5
68.0
53.6
59.6
Table 2: RoboCasa-GR1 Tabletop Tasks. The 24 tasks are partitioned into five evaluation groups. Each entry is the unweighted mean success rate (%) within the corresponding group (Overall averages over all 24 tasks). Best result in each column is shown in bold .
Figure 4: Real-robot setup and evaluation conditions. (a) AgileX Cobot Magic, an ALOHA-style platform; only the right PiPER arm is used. (b) Nominal and shifted conditions involving background, manipulation height, and objects for evaluating the frozen policy. (c) Gaussian observation noise and dynamic lighting conditions for test-time training, evaluated separately from the frozen-policy results.
Table 6
Fusion
MoT
Lalign
Ldyn
Avg.
✓
✓
✓
53.9 ( ↓ 14.6)
✓
✓
✓
65.4 ( ↓ 3.1)
✓
✓
✓
60.8 ( ↓ 7.7)
✓
✓
60.4 ( ↓ 8.1)
✓
✓
✓
65.6 ( ↓ 2.9)
✓
✓
✓
✓
68.5
Table 5: Component ablation on SimplerEnv (average success, %). Each row removes one component; dark red arrows give the drop from the full model. Evaluation follows the main SimplerEnv protocol.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Spatial feature structure. Final-layer attention (left) and PCA of patch embeddings (right) for V-JEPA 2, LeWM, and our control-aligned encoder trained with the Dynamic CLS loss (Ours). LeWM focuses on the manipulation region, while Ours shows more concentrated attention along robot and object boundaries than LeWM. PCA is fitted separately on the same frames; colors are comparable only within each encoder.
Figure 6: CLS transition similarity across action patterns. (a) Pairwise cosine similarity between CLS transition vectors Δct=ct+1−ct ; warmer colors indicate higher similarity. The green block (A) covers rightward-translation transitions 1→2 through 3→4 , and the yellow block (B) covers gripper-closing transitions 6→7 through 9→10 . The black-outlined cell (C) contrasts 1→2 with 7→8 , illustrating low similarity across the two action patterns. (b) Corresponding observation sequences and the cross-action comparison; dashed white circles highlight the gripper region.
Dataset
Episodes
Frames
Rate (Hz)
Pretraining windows
BridgeData V2
53,192
1,893,026
5
1,520,923
Fractal
87,212
3,786,400
3
3,178,323
Appendix
Table 6: Bridge/Fractal data inventory. Encoder pretraining windows use four observations at stride two and do not cross episode boundaries.
Parameter
Bridge/Fractal
RoboCasa-GR1
PiPER fine-tuning
Update budget
100,000
100,000
50,000
Devices × batch per device
8×16
8×8
8×16
Global batch size
128
64
128
Backbone learning rate
10−5
10−5
10−4 (LoRA)
Predictive-branch learning rate
10−4
10−4
10−4 (LoRA)
Action-decoder learning rate
10−4
10−4
10−5
Appendix
Table 7: Policy optimization. Shared optimizer, precision, and backbone settings are described in the text.
Figure 7: Real-robot workstation with additional colored illumination. This lighting is used for dynamic-lighting test-time training, not for the frozen-policy columns in Table 4 .
Vision-Language-Action (VLA) models commonly adapt pretrained Vision-Language Models (VLMs) to robot control by mapping visual observations and language instructions to continuous actions. Existing approaches typically take an action-insufficiency view, assuming that pretrained VLM latents either lack directly usable action information or should be shielded from action-learning signals. Against this view, our \textit{Quotient Theory for VLA} shows that pretrained VLM latents are not action-insufficient but action-sufficient: they already contain the information needed for control, yet remain overcomplete by distinguishing prompt-level variations that induce the same optimal action behavior. To operationalize this theory, we propose QuoVLA, a quotient-space framework for VLA that compresses pretrained VLM latents into action-sufficient representations. Specifically, QuoVLA instantiates this principle with a quantization module and a dual-branch design with relative temporal-complexity regularization, preserving action-relevant information while removing prompt-level redundancy. Extensive experiments across multiple benchmarks demonstrate that QuoVLA achieves strong performance, with particularly notable improvements in generalization under visual, linguistic, and environmental distribution shifts. Our code will be made publicly available.
Learning to predict how the world evolves can provide vision-language-action (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation. We introduce PLaW-VLA, which models task-relevant future states in a pretrained prediction-oriented representation space, reducing the need to predict control-irrelevant visual details. Built on a Mixture-of-Transformers architecture, PLaW-VLA conditions action generation on observation history, current task semantics, and predicted future states through structured causal attention. Experiments show a +11.8 percentage-point (pp) gain over reactive policies on RoboTwin Hard Horizon III and a +1.77 pp gain over reconstruction-oriented latent prediction on zero-shot LIBERO-Plus, supporting improved long-horizon control and generalization under distribution shift, respectively. By avoiding low-level visual reconstruction, PLaW-VLA lowers the burden of future prediction, enabling a lightweight latent world model with parallel future prediction and about 1/19 the inference latency of generative world-action modeling at comparable policy performance.
Yu Liu, Hetian Guo, Tianlv Huang +11
Jilin University · Astribot · Harbin Institute of Technology, Shenzhen +2
Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence. Through the shared predictor, transition supervision directly shapes the backbone, from which dedicated representations are extracted for action prediction. The same design can also be instantiated in pretrained VLA policies while preserving their original perception and action pathways. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained π0.5 instantiation reaches 86.3%, achieving the best overall performance. Experiments on RoboTwin 2.0 and real-world bimanual manipulation further demonstrate strong generalization under visual and spatial shifts.
Yihan Lin, Jiawei He, Shifeng Bao +6
School of Information, Renmin University of China, Beijing, China · XYZ Embodied AI, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China +3