RoboJEPA: Scaling Robotic Latent World Models
Organizations: FAIR at Meta · Chandar Research Lab · Mila - Quebec AI Institute · Polytechnique Montréal · Work done at Meta
Abstract
Latent world models have shown a remarkable ability to predict future states and to plan in the real world. In practice, however, we lack a principled way to estimate how their capabilities scale with model size, data, and compute, an open problem that slows progress in the field. In this work we present RoboJEPA, a world model based on the Joint Embedding Predictive Architecture (JEPA) and trained on a large-scale dataset spanning 12 robotic embodiments. We show that RoboJEPA's imagination error, the error of its latent rollouts, follows a second-order power law in compute, allowing us to predict model quality well beyond the scale at which the law is fit. We further show that downstream robotic planning performance improves predictably with compute, and that imagination error is strongly correlated with it, making it a reliable proxy for real-robot evaluation. Finally, we demonstrate that latent world models can be deployed zero-shot as robotic agents, planning toward a single goal image to solve tasks requiring long-horizon planning on real hardware. We release all model checkpoints together with our training and robot deployment code. To our knowledge, this is the first work to establish scaling laws for multi-embodiment robotic world models trained on real robot data, and RoboJEPA, at 8B parameters, is the largest JEPA predictor model trained to date.
Figures & tables
| Dataset group | Episodes | Video hours | Action hours | Embodiments |
|---|---|---|---|---|
| DROID ( Khazatsky et al., 2024 ) | 92k | 414 | 138 | Franka |
| Roboset ( Bharadhwaj et al., 2024 ) | 73k | 1244 | 380 | Franka (Roboset) |
| RoboMind ( Wu et al., 2024b ) | 35.3k | 276 | 98 | Franka, AgileX |
| LeRobot ( Cadene et al., 2024 ) | 21.5k | 182 | 93 | SO-101 |
| RoboCasa365 ( Nasiriany et al., 2026 ) | 32k | 482 | 482 | Franka + Omron (sim) |
| 1X ( 1X Technologies, 2024 ) | 23.5k | 90 | 90 | 1X humanoid |
| Parametric Curve | Curve parameters | Extrapolation Error (4B and 8B), | ||
| DROID | RoboCasa | DROID | RoboCasa | |
| BNSL ( Caballero et al., 2022 ) , | Section J.1 | |||
| UNSL ( Caballero et al., 2026 ) , | Section J.2 | |||
| Model | Goal | Grasp | Object Lift | Pick and Place | |||
|---|---|---|---|---|---|---|---|
| Progress | Success | Progress | Success | Progress | Success | ||
| -FAST | Text | 35 | 22 | 20 | 12 | 55 | 40 |
| 25 | 5 | 12 | 0 | 69 | 53 | ||
| RoboJEPA-4B | Image | 60 | 60 | 54 | 30 | 49 | 21 |
| RoboJEPA-8B | 67 | 67 | 65 | 50 | 42 | 27 | |
Appendix figures & tables41 assets
Supplementary material from the paper’s appendix.
Appendix
| V-JEPA 2-AC ( Assran et al., 2025 ) | RoboJEPA (ours) | |
| Architecture | ||
| Base ViT architecture | V-JEPA 2 | |
| Training parameterisation | Feed-forward next-embedding (frame) prediction with frame-causal mask; trained with an latent loss. | |
| Normalisation | LayerNorm | RMSNorm |
| QK-norm | No | Yes |
| Position embedding | 3-D spatio-temporal RoPE | 4-D RoPE (space–time views) |
| # | Platform | Notes |
|---|---|---|
| default state-delta ids (one per physical platform) | ||
| 1 | Franka | DROID, RoboMind-Franka, Roboset, OXE-FMB (default) |
| 2 | Franka + Omron | RoboCasa (action-key) |
| 3 | SO-101 | LeRobot |
| 4 | AgiBot dual-arm (Robotiq) | AgiBot World |
| 5 | AgiBot humanoid | AgiBot World |
| Hyperparameter | 22M | 50M | 100M | 300M | 1B | 2B | 4B | 8B | Notation |
| Predictor depth | |||||||||
| Predictor width | |||||||||
| Query heads | |||||||||
| KV heads (GQA) | |||||||||
| Head dim | |||||||||
| MLP hidden ratio |
| Dataset | Traj. | Vid. hrs | Act. hrs | Resolution | FPS | Embodiment |
| DROID ( Khazatsky et al., 2024 ) | 92k | 414 | 138 | 60 | Franka | |
| Roboset ( Bharadhwaj et al., 2024 ) | 73k | 1244 | 380 | 5 | Franka (Roboset) | |
| RoboMind Franka ( Wu et al., 2024b ) | 25k | 92 | 37 | 30 | Franka | |
| RoboMind AgileX ( Wu et al., 2024b ) | 10.3k | 184 | 61 | 30 | AgileX | |
| LeRobot ( Cadene et al., 2024 ) | 21.5k | 182 | 93 | 30 | SO-101 | |
| RoboCasa365 ( Nasiriany et al., 2026 ) | 32k | 482 | 482 | 20 | Franka + Omron (sim) |
| Dataset (action-source) | 1-view weight | 2-view weight | Total | Marginal dataset weight |
|---|---|---|---|---|
| DROID (state-delta cartesian) | 0.1692 | |||
| DROID (state-delta joints) | ||||
| DROID (action-key cartesian) | ||||
| DROID (action-key joints) | ||||
| DROID (action-key cartesian-vel) | ||||
| DROID (action-key joint-vel) |
| Semantic id | |||||||
|---|---|---|---|---|---|---|---|
| Role | Left exterior | Right exterior | Wrist / eye-in-hand | Head / front-facing | Default / unmapped | Top-down | Planar-arm overhead |
| DROID | left | right | wrist | — | — | — | — |
| Roboset | left | right | wrist | — | — | top | — |
| RoboMind | camera_left | camera_right | — | camera_front | — | camera_top | — |
| RoboMind AgileX | — | — | camera_{left,right}_wrist | camera_front | — | — | — |
| RoboCasa | robot0_agentview_left | robot0_agentview_right | robot0_eye_in_hand | — | — | — | — |
| Hyperparameter | Symbol | Value |
| Loss (defined once in Equation 5 ) | ||
| Pre-loss normalization | parameter-free LayerNorm on | |
| Rollout step index | — | |
| Rollout prefix sampling | uniform on | |
| Recurrent rollout input | — | detached from autograd graph |
| Rollout / TF mixing | — | |
| Hyperparameter | Symbol | Value |
| Architecture | ||
| DiT depth | 32 | |
| Hidden dim | 1280 | |
| Query / KV heads | – | 20 / 10 (GQA, ( Ainslie et al., 2023 ) ) |
| MLP ratio | – | 4.0 |
| Cosmos latent channels | – | 16 |
| Model | Goal specification | Grasp | Object Lift | Pick and Place | |||
|---|---|---|---|---|---|---|---|
| Task Progress | Success Rate | Task Progress | Success Rate | Task Progress | Success Rate | ||
| Model finetuned on DROID, no in-domain finetuning | |||||||
| -FAST | text instruction | ||||||
| Pretrained models, no DROID finetuning, no in-domain finetuning, a unified planner for all tasks | |||||||
| RoboJEPA-22M | single goal image | ||||||
| Hyperparameter | Symbol | Value |
| Horizon and closed loop | ||
| Planning horizon (predictor steps) | ||
| Context length (encoder steps) | ||
| Execution prefix (actions per replan) | ||
| Environment steps per episode | — | |
| CEM search | ||
| Symbol | Reach | ObstacleReach | ObjectReach | Push | |
|---|---|---|---|---|---|
| Planning hyperparameters | |||||
| Predictor horizon | |||||
| CEM samples per iteration (total) | — | ||||
| CEM iterations (cold / warm) | / | ||||
| Elite count, top- | |||||
| Computational cost | per-token in space | ||||
| Configuration | Wall (ms) | GPU kernel (ms) | Dispatch (ms) | Dispatch % |
|---|---|---|---|---|
| 1-view eager | 145.6 | 76.7 | 68.9 | 47.3% |
| 1-view AOT | 45.0 | 41.9 | 3.1 | 6.8% |
| 2-view eager | 217.1 | 148.7 | 68.4 | 31.5% |
| 2-view AOT | 96.7 | 93.7 | 3.0 | 3.1% |
| Cache step | BF16 AOTI (ms) | FP8 torch.compile (ms) | Speedup |
| 61.5 | 54.8 | ||
| 94.2 | 82.4 | ||
| – | |||
| average | 93.1 | 84.1 |
| Dataset | ||||||
|---|---|---|---|---|---|---|
| DROID | ||||||
| RoboCasa |
| Dataset | Fitted parameters |
|---|---|
| DROID | |
| RoboCasa |
| Model size | Short-horizon training | Long-horizon training |
|---|---|---|
| ( , ) | ( , ) | |
| M | ||
| M | ||
| M | ||
| M | ||
| B |
| Stage | Recipe | Total per-step PFLOPs |
|---|---|---|
| Pretraining / short-horizon training | , ( ) | PF |
| Long-horizon training | with KV cache, multi-res ( ) | PF |
| DROID finetune | -view , with KV cache ( ) | PF |
| Type | Benchmark | Metric | V-JEPA 2 | V-JEPA 2.1 |
| ViT-g (1B) | ViT-G (2B) | |||
| Global | SSv2 action recognition | Top-1 (%) | 77.3 | 77.7 |
| Global | K400 action recognition | Top-1 (%) | 87.3 | 87.7 |
| Global | IN1K image classification | Top-1 (%) | 85.1 | 85.5 |
| Global | IntPhys2 ( Bordes et al., 2025 ) | VOE (%) | 66.0 | 58.8 |
| Dense | ADE20K semantic segmentation | mIoU (%) | 24.4 | 47.9 |