Organizations: University of Science and Technology of China · Eastern Institute of Technology, Ningbo · Hangzhou Dianzi University · ShanghaiTech University
Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from 60.9% to 68.5% over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches 72.7%; on a real robot, it retains 70%--75% success under background, height, and object shifts where the base policy collapses to 0%.
Figures & tables
Figure 1: We present Juno, a unified framework that makes predictive latents useful across the vision-language-action (VLA) lifecycle. Juno integrates control-aligned JEPA pretraining, visual feature fusion and decoupled future-latent distillation, and dynamics-first test-time adaptation that updates the world model before re-aligning the policy. It improves performance on SimplerEnv and RoboCasa and supports robust real-world manipulation under distribution shifts.
Figure 2: (a) Patch-wise temporal cosine distance between frames t⋆ and t⋆+Δt for LeWM and V-JEPA2. (b) Temporal state similarity cos(ct,ct+Δt) over the episode, spanning slight motion ( t≤14 ) and pick-and-place ( t>14 ). LeWM (CLS) stays nearly constant, whereas LeWM+ Ldyn becomes sensitive during pick-and-place; mean-pooled baselines are shown for reference.
Figure 3: Overview of Juno. Top left: Dynamic CLS pretraining aligns CLS transitions with motion-weighted patch changes from a frozen JEPA supervisor. Bottom left: JEPA features enrich visual tokens, while a decoupled predictive branch distills future latent states to jointly guide action generation. Right: Test-time adaptation updates the world model using both successful and failed executions, then re-aligns the policy on verified successful executions. Flame and snowflake icons indicate trainable and frozen modules, respectively.
Method
PnP + Close
PnP From Cuttingboard
PnP From Placemat
PnP From Plate
PnP From Tray
Avg.
GR00T-N1.6
24.2
56.9
51.9
57.6
55.1
47.6
Qwen3PI
42.3
46.0
43.5
44.0
44.0
43.9
Qwen3OFT
43.7
50.4
41.5
61.0
49.2
48.8
Qwen3FAST
35.0
50.4
33.5
45.0
32.0
39.0
Qwen3GR00T
50.3
52.8
38.0
58.5
39.2
47.8
Juno (Ours)
57.3
64.8
55.5
68.0
53.6
59.6
Table 2: RoboCasa-GR1 Tabletop Tasks. The 24 tasks are partitioned into five evaluation groups. Each entry is the unweighted mean success rate (%) within the corresponding group (Overall averages over all 24 tasks). Best result in each column is shown in bold .
Figure 4: Real-robot setup and evaluation conditions. (a) AgileX Cobot Magic, an ALOHA-style platform; only the right PiPER arm is used. (b) Nominal and shifted conditions involving background, manipulation height, and objects for evaluating the frozen policy. (c) Gaussian observation noise and dynamic lighting conditions for test-time training, evaluated separately from the frozen-policy results.
Table 6
Fusion
MoT
Lalign
Ldyn
Avg.
✓
✓
✓
53.9 ( ↓ 14.6)
✓
✓
✓
65.4 ( ↓ 3.1)
✓
✓
✓
60.8 ( ↓ 7.7)
✓
✓
60.4 ( ↓ 8.1)
✓
✓
✓
65.6 ( ↓ 2.9)
✓
✓
✓
✓
68.5
Table 5: Component ablation on SimplerEnv (average success, %). Each row removes one component; dark red arrows give the drop from the full model. Evaluation follows the main SimplerEnv protocol.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Spatial feature structure. Final-layer attention (left) and PCA of patch embeddings (right) for V-JEPA 2, LeWM, and our control-aligned encoder trained with the Dynamic CLS loss (Ours). LeWM focuses on the manipulation region, while Ours shows more concentrated attention along robot and object boundaries than LeWM. PCA is fitted separately on the same frames; colors are comparable only within each encoder.
Figure 6: CLS transition similarity across action patterns. (a) Pairwise cosine similarity between CLS transition vectors Δct=ct+1−ct ; warmer colors indicate higher similarity. The green block (A) covers rightward-translation transitions 1→2 through 3→4 , and the yellow block (B) covers gripper-closing transitions 6→7 through 9→10 . The black-outlined cell (C) contrasts 1→2 with 7→8 , illustrating low similarity across the two action patterns. (b) Corresponding observation sequences and the cross-action comparison; dashed white circles highlight the gripper region.
Dataset
Episodes
Frames
Rate (Hz)
Pretraining windows
BridgeData V2
53,192
1,893,026
5
1,520,923
Fractal
87,212
3,786,400
3
3,178,323
Appendix
Table 6: Bridge/Fractal data inventory. Encoder pretraining windows use four observations at stride two and do not cross episode boundaries.
Parameter
Bridge/Fractal
RoboCasa-GR1
PiPER fine-tuning
Update budget
100,000
100,000
50,000
Devices × batch per device
8×16
8×8
8×16
Global batch size
128
64
128
Backbone learning rate
10−5
10−5
10−4 (LoRA)
Predictive-branch learning rate
10−4
10−4
10−4 (LoRA)
Action-decoder learning rate
10−4
10−4
10−5
Appendix
Table 7: Policy optimization. Shared optimizer, precision, and backbone settings are described in the text.
Figure 7: Real-robot workstation with additional colored illumination. This lighting is used for dynamic-lighting test-time training, not for the frozen-policy columns in Table 4 .
School of Information, Renmin University of China, Beijing, China · XYZ Embodied AI, Beijing, China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China +3