Organizations: University of Science and Technology of China · Eastern Institute of Technology, Ningbo · Hangzhou Dianzi University · ShanghaiTech University
Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from 60.9% to 68.5% over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches 72.7%; on a real robot, it retains 70%--75% success under background, height, and object shifts where the base policy collapses to 0%.
Figures & tables
Figure 1: We present Juno, a unified framework that makes predictive latents useful across the vision-language-action (VLA) lifecycle. Juno integrates control-aligned JEPA pretraining, visual feature fusion and decoupled future-latent distillation, and dynamics-first test-time adaptation that updates the world model before re-aligning the policy. It improves performance on SimplerEnv and RoboCasa and supports robust real-world manipulation under distribution shifts.
Figure 2: (a) Patch-wise temporal cosine distance between frames t⋆ and t⋆+Δt for LeWM and V-JEPA2. (b) Temporal state similarity cos(ct,ct+Δt) over the episode, spanning slight motion ( t≤14 ) and pick-and-place ( t>14 ). LeWM (CLS) stays nearly constant, whereas LeWM+ Ldyn becomes sensitive during pick-and-place; mean-pooled baselines are shown for reference.
Figure 3: Overview of Juno. Top left: Dynamic CLS pretraining aligns CLS transitions with motion-weighted patch changes from a frozen JEPA supervisor. Bottom left: JEPA features enrich visual tokens, while a decoupled predictive branch distills future latent states to jointly guide action generation. Right: Test-time adaptation updates the world model using both successful and failed executions, then re-aligns the policy on verified successful executions. Flame and snowflake icons indicate trainable and frozen modules, respectively.
Method
PnP + Close
PnP From Cuttingboard
PnP From Placemat
PnP From Plate
PnP From Tray
Avg.
GR00T-N1.6
24.2
56.9
51.9
57.6
55.1
47.6
Qwen3PI
42.3
46.0
43.5
44.0
44.0
43.9
Qwen3OFT
43.7
50.4
41.5
61.0
49.2
48.8
Qwen3FAST
35.0
50.4
33.5
45.0
32.0
39.0
Qwen3GR00T
50.3
52.8
38.0
58.5
39.2
47.8
Juno (Ours)
57.3
64.8
55.5
68.0
53.6
59.6
Table 2: RoboCasa-GR1 Tabletop Tasks. The 24 tasks are partitioned into five evaluation groups. Each entry is the unweighted mean success rate (%) within the corresponding group (Overall averages over all 24 tasks). Best result in each column is shown in bold .
Figure 4: Real-robot setup and evaluation conditions. (a) AgileX Cobot Magic, an ALOHA-style platform; only the right PiPER arm is used. (b) Nominal and shifted conditions involving background, manipulation height, and objects for evaluating the frozen policy. (c) Gaussian observation noise and dynamic lighting conditions for test-time training, evaluated separately from the frozen-policy results.
Table 6
Fusion
MoT
Lalign
Ldyn
Avg.
✓
✓
✓
53.9 ( ↓ 14.6)
✓
✓
✓
65.4 ( ↓ 3.1)
✓
✓
✓
60.8 ( ↓ 7.7)
✓
✓
60.4 ( ↓ 8.1)
✓
✓
✓
65.6 ( ↓ 2.9)
✓
✓
✓
✓
68.5
Table 5: Component ablation on SimplerEnv (average success, %). Each row removes one component; dark red arrows give the drop from the full model. Evaluation follows the main SimplerEnv protocol.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Spatial feature structure. Final-layer attention (left) and PCA of patch embeddings (right) for V-JEPA 2, LeWM, and our control-aligned encoder trained with the Dynamic CLS loss (Ours). LeWM focuses on the manipulation region, while Ours shows more concentrated attention along robot and object boundaries than LeWM. PCA is fitted separately on the same frames; colors are comparable only within each encoder.
Figure 6: CLS transition similarity across action patterns. (a) Pairwise cosine similarity between CLS transition vectors Δct=ct+1−ct ; warmer colors indicate higher similarity. The green block (A) covers rightward-translation transitions 1→2 through 3→4 , and the yellow block (B) covers gripper-closing transitions 6→7 through 9→10 . The black-outlined cell (C) contrasts 1→2 with 7→8 , illustrating low similarity across the two action patterns. (b) Corresponding observation sequences and the cross-action comparison; dashed white circles highlight the gripper region.
Dataset
Episodes
Frames
Rate (Hz)
Pretraining windows
BridgeData V2
53,192
1,893,026
5
1,520,923
Fractal
87,212
3,786,400
3
3,178,323
Appendix
Table 6: Bridge/Fractal data inventory. Encoder pretraining windows use four observations at stride two and do not cross episode boundaries.
Parameter
Bridge/Fractal
RoboCasa-GR1
PiPER fine-tuning
Update budget
100,000
100,000
50,000
Devices × batch per device
8×16
8×8
8×16
Global batch size
128
64
128
Backbone learning rate
10−5
10−5
10−4 (LoRA)
Predictive-branch learning rate
10−4
10−4
10−4 (LoRA)
Action-decoder learning rate
10−4
10−4
10−5
Appendix
Table 7: Policy optimization. Shared optimizer, precision, and backbone settings are described in the text.
Figure 7: Real-robot workstation with additional colored illumination. This lighting is used for dynamic-lighting test-time training, not for the frozen-policy columns in Table 4 .
Vision-Language-Action (VLA) models commonly adapt pretrained Vision-Language Models (VLMs) to robot control by mapping visual observations and language instructions to continuous actions. Existing approaches typically take an action-insufficiency view, assuming that pretrained VLM latents either lack directly usable action information or should be shielded from action-learning signals. Against this view, our \textit{Quotient Theory for VLA} shows that pretrained VLM latents are not action-insufficient but action-sufficient: they already contain the information needed for control, yet remain overcomplete by distinguishing prompt-level variations that induce the same optimal action behavior. To operationalize this theory, we propose QuoVLA, a quotient-space framework for VLA that compresses pretrained VLM latents into action-sufficient representations. Specifically, QuoVLA instantiates this principle with a quantization module and a dual-branch design with relative temporal-complexity regularization, preserving action-relevant information while removing prompt-level redundancy. Extensive experiments across multiple benchmarks demonstrate that QuoVLA achieves strong performance, with particularly notable improvements in generalization under visual, linguistic, and environmental distribution shifts. Our code will be made publicly available.
Robust robot control benefits from explicitly modeling state transitions, but video-generation world action models (WAMs) introduce substantial deployment cost. Existing latent WAMs avoid explicit future generation, but often compress predictive representations or separate predictive modeling from the representations used for action generation. We introduce JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor. JEPA-WAM predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, while preserving dense patch-level correspondence. Through the shared predictor, transition supervision directly shapes the backbone, from which dedicated representations are extracted for action prediction. The same design can also be instantiated in pretrained VLA policies while preserving their original perception and action pathways. On LIBERO-Plus, JEPA-WAM achieves 79.2%, the best result without large-scale robot-policy pretraining, while its pretrained π0.5 instantiation reaches 86.3%, achieving the best overall performance. Experiments on RoboTwin 2.0 and real-world bimanual manipulation further demonstrate strong generalization under visual and spatial shifts.
Yihan Lin, Jiawei He, Shifeng Bao +6
School of Information, Renmin University of China, Beijing, China · XYZ Embodied AI, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China +3
World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned from scratch in a single downstream stage, with the predictor's future-informed context key--value states conditioning action generation. With 0.9B total parameters, of which 0.6B are trainable, V-JEPA Policy achieves competitive performance with representative WAM and vision-language-action baselines across LIBERO, LIBERO-Plus, and RoboCasa-GR1. Comparing visual foundations under the same downstream framework and training budget identifies V-JEPA latents as more effective than the discriminative, reconstructive, and video-understanding-oriented alternatives, particularly under distribution shifts. Beyond task-specific learning, pretraining the predictor on DROID video--instruction pairs without action labels and adapting it into a WAM yields substantial gains in downstream control and out-of-distribution generalization. Together, these findings establish predictive visual latents as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos. Our code is available at https://github.com/breez3young/VJEPA-Policy.
Yang Zhang, Jiangyuan Zhao, Chenyou Fan +4
Tsinghua University · Shanghai Jiao Tong University · Fudan University +4