cs.ROOct 7, 2026

Juno: Taming Predictive Latents for Vision-Language-Action Models

Authors: Yuchen Zhu, Chenyi Xu, Yulin Zhang, Gang Xu, Wentao Zhu

Organizations: University of Science and Technology of China · Eastern Institute of Technology, Ningbo · Hangzhou Dianzi University · ShanghaiTech University

Abstract

Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from 60.9%60.9\% to 68.5%68.5\% over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches 72.7%72.7\%; on a real robot, it retains 70%70\%--75%75\% success under background, height, and object shifts where the base policy collapses to 0%0\%.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. QuoVLA: Quotient Space for Vision-Language-Action Models

    May 24, 2026Xuan Wang, Yinan Wu, Haoran Duan +1Video Latents

  2. JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

    Aug 10, 2026Yihan Lin, Jiawei He, Shifeng Bao +6Efficient World-Action ModelRobot Policies

  3. V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents

    Sep 29, 2026Yang Zhang, Jiangyuan Zhao, Chenyou Fan +4Video Joint Embedding Predictive ArchitectureRecent World-Action Models