Robotic systems often exhibit unstable modes, along which small perturbations and disturbances can cause unbounded growth unless corrected through feedback. Controlling such systems from high-dimensional visual observations requires representations that preserve these modes. Joint-embedding predictive architectures (JEPAs) provide a natural framework for learning such representations and their dynamics from visual data. However, we demonstrate that next step prediction combined with anti-collapse regularization does not guarantee that controllable unstable modes are preserved: the training loss can be minimized while these modes are collapsed, making stabilization from the learned representation impossible. To address this, we augment world-model training with an action reconstruction objective (i.e., an inverse dynamics loss) that encourages control-aware representations, namely, visual representations that preserve crucial features for control. We prove that exact action reconstruction makes the encoder injective on the finite-horizon reachable subspace. Thus, the encoder cannot discard any state direction reachable by an action sequence within H steps. Moreover, we show that, as H grows, the dominant eigenspace of the finite-horizon controllability Gramian converges to the controllable unstable subspace. We establish our theoretical results for linear systems and demonstrate empirically that our findings extend to nonlinear visual control tasks (CartPole, Walker2D, and PointMaze), highlighting the benefits of control-aware representation learning.
Figures & tables
Figure 1 : Architecture of our proposed world model. At each step, an encoder ( E ) maps observation yt to a latent state zt , and a predictor ( P ) rolls out latent predictions z^t+1,…,z^t+H conditioned on actions. Training losses are applied at each predicted latent: 1SP/MSP (single- or multi-step prediction) aligns z^t+k with the encoded ground-truth zt+k . The EP-IDM estimates the intermediate actions a^t,H=(a^t,a^t+1,…,a^t+H−1) from zt and zt+H .
Figure 2 : SIG vs. EP-IDM across two examples. Each row compares a representation trained with SIGReg (SIG) against one trained with the endpoint inverse-dynamics (action-reconstruction) loss (EP-IDM). Top row: synthetic open-loop unstable LDS. Bottom row: linearized CartPole. Left: true unstable ( vu , green) and stable ( vs , grey) eigenvectors versus the learned predictor’s dominant eigenvector, for SIG (blue) and EP-IDM (dashed orange). Center: closed-loop trajectories under the latent LQR controller from four initial states (circles) for the model trained with SIGReg. Right: closed-loop trajectories under the latent LQR controller for the model trained with the inverse dynamics loss (EP-IDM). For the bottom row x3 , and x4 corresponds to the pole angle and angular velocity, respectively. The latent dimension is one for the example in the top row and three for the example in the bottom row.
Figure 3 : CartPole , Walker2D , and PointMaze visual-control tasks used in our evaluation. Each row depicts ten frames from a successful trial (initial state: blue border, final state: green border). CartPole: a continuous-action inverted-pendulum balancing task controlled via LQR. Walker2D: a bipedal robot locomotion task controlled via iCEM ( Pinneri et al., 2021 ) . PointMaze: a 2-D point-mass robot navigation task planned via CEM.
Figure 4 : CartPole state norm ∥xt∥2 under the designed latent LQR controller over 300 steps.
Figure 5 : Walker2D. Phase portraits of right-hip angle vs. angular velocity for GT dynamics, 1SP+EP-IDM, 1SP+SIG, and MSP+SIG. Each trajectory is a 500 -step closed-loop rollout: GT uses real MuJoCo physics, model panels roll out in latent space under an MLP-decoded state, with a SAC ( Huang et al., 2022 ; Haarnoja et al., 2018 ) policy acting on the decoded observation.
Figure 6 : Trajectory frames for Walker2D . Each row shows ten frames uniformly sampled from a 500 -step closed-loop rollout under the SAC policy ( Haarnoja et al., 2018 ) , decoded from the trained model’s latent state via an MLP decoder.
Model
Avg Velocity (m/s)
Displacement (m)
Final Height (m)
Ground truth (GT)
3.61
14.43
1.21
1SP + SIG
0.08
0.06
0.98
MSP + SIG
0.46
0.66
0.85
1SP + EP-IDM
2.91
9.15
0.92
Table 2 : Walker2D locomotion under latent iCEM ( Pinneri et al., 2021 ) . All metrics are averaged over ten trials. Final height is the torso height (m) at the end of the episode.
Figure 7 : Planning trajectory frames for Walker2D . Each row shows ten frames uniformly sampled from a 500 -step closed-loop rollout under latent iCEM planning for the trained models. We select the best trial, with respect to the forward displacement, across the ten trials for the trained models.
Figure 8 : CartPole diagnostics. Left: Phase portraits of ground-truth and learned vector fields. Center: H=3 open-loop prediction error across (θ,θ˙) . Right: Zero-action planning cost.
Encoder
Objective
CEM (SR)
LQR (SR)
GBP (SR)
GT
100%
80%
90.0% *
From scratch
1SP + IDM
100%
70%
100%
MSP + MS-IDM
100%
80%
90.0%
MSP + EP-IDM
90.0%
80%
90.0%
1SP + EP-IDM
100%
80%
80.0%
MSP + SIG
90.0%
60%
100%
Table 3 : PointMaze navigation success rate (%).
Figure 9 : Local stability on CartPole . First two panels: Empirical region of attraction of the learned LQR controller when deployed in the ground-truth system. The dashed line indicates the ground-truth LQR boundary. Last two panels: Lyapunov decrease ΔV=V(x′)−V(x) under the latent LQR controller.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10 : CartPole trajectory visualization for the trained models under the latent LQR controller.
Figure 11 : MSP+EP-IDM+SIG model on CartPole . Left: Phase portrait of ground-truth vs. learned vector fields. Center: H=3 open-loop prediction error. Right: Zero-action planning cost.
Figure 12 : Local stability of MSP+EP-IDM+SIG on CartPole . Left: Empirical region of attraction (ROA) of the learned LQR policy in the GT environment. Dashed line: GT LQR boundary. Right: Lyapunov decrease ΔV=V(x′)−V(x) under the learned policy using the GT DARE solution.
Figure 13 : MSP+IBOT+PR-EP-IDM model on CartPole . Panel descriptions are in Fig. 11 .
Figure 14 : Local stability of MSP+IBOT+PR-EP-IDM on CartPole . Panel descriptions are the same as those in Fig. 12 .
Figure 15 : MSP+IBOT+PR-SIG model on CartPole . Panel descriptions are the same as those in Fig. 11 .
Figure 16 : MSP+IBOT model on CartPole . Panel descriptions are the same as those in Fig. 11 .
Figure 17 : Local stability of MSP+IBOT on CartPole . Left: Empirical region of attraction (ROA) of the learned LQR policy in the GT environment. Panel descriptions are as those in Fig. 12 .
Figure 18 : DINO-WM model on CartPole . Panel descriptions are the same as those in Fig. 11 .
Encoder
Objective
CEM ( 1/L2 )
CEM (FE)
GT
0.0003
0.30202
From scratch
1SP + SIG
0.0003
0.19956
1SP + IDM + MS-IDM
0.0003
0.43608
1SP + EP-IDM
0.0003
0.13683
MSP + SIG
0.0003
0.07667
MSP + EP-IDM + SIG
0.0003
0.11985
Appendix
Table 4 : CartPole CEM diagnostic metrics (action context 1). 1/L2 : inverse squared mean stabilization length. FE: final error ∥xT∥ . PR denotes a learned projector on top of the frozen encoder.
Figure 19 : PointMaze visualization for our trained models under CEM planning.
Figure 20 : Latent goal-distance fields ∥z−zg∥ across the PointMaze U-maze for four trained world models.
Figure 21 : k -step closed-loop prediction error along trajectories generated by the ground-truth LQR controller. Curves show the physical-state error between the ground-truth trajectory and the trajectory decoded from each latent rollout.
Encoder
Objective
LQR (SR)
LQR (MFS)
CEM (SR)
GBP (SR)
DINOv2
1SP
0%
0.301
0%
0%
MSP
0%
0.232
0%
0%
1SP + PR-EP-IDM
0%
0.025
20%
0%
MSP + PR-SIG
0%
0.460
100%
0%
MSP + PR-IDM
0%
0.455
0%
0%
MSP + PR-EP-IDM
0%
0.513
40%
10%
Appendix
Table 5 : CartPole control performance with frozen pretrained encoders. SR: success rate. MFS: mean fraction of steps within the success threshold. PR denotes a learned projector. DINO-WM ( Zhou et al., 2024 ) is an external baseline.
Figure 22 : CartPole state norm ∥xt∥2 under the learned latent LQR controller over 300 steps. Left: DINOv2. Right: iBOT.
Task
Image size
Proprioception
Action
Frame skip
Training trajectories
CartPole
128×128
4
1
5
787
Walker2D
64×64
17
6
5
3,200
PointMaze
64×64
4
2
5
1,600
Appendix
Table 6 : Task-specific observation and action dimensions. The frame skip is the number of simulator steps represented by one model transition.
Hyperparameter
Value
Hyperparameter
Value
Optimizer
AdamW
Epochs
200
Batch size
64
Prediction horizon H
3
Initial learning rate
10−4
Minimum learning rate
10−6
Weight decay
10−3
Warm-up epochs
5
Gradient-norm threshold
1
Latent dimension
192
Random seed
42
Active loss weights
1
Appendix
Table 7 : World-model training hyperparameters.
Task
Planner
Horizon
Execute
Population
Elites
Iterations
CartPole
CEM
10
1
300
30
10
Walker2D
iCEM
3
1
200
20
5
PointMaze
CEM
25
25
300
30
10
Appendix
Table 8 : CEM and improved-CEM planning hyperparameters. Here, “execute” denotes the number of actions applied before replanning.
Objective
LQR (SR)
LQR (MFS)
CEM (SR)
GBP (SR)
1SP + IDM + MS-IDM
100%
0.999
100%
100%
1SP + EP-IDM
100%
0.999
100%
100%
MSP + EP-IDM + SIG
100%
0.996
100%
100%
Appendix
Table 9 : CartPole ablation over inverse-dynamics formulations for an encoder trained from scratch.
LQR (SR)
CEM (SR)
GBP (SR)
Encoder
Objective
Context 1
Context 5
Context 1
Context 5
Context 1
Context 5
From scratch
1SP + SIG
0
0
100
100
90
90
1SP + IDM + MS-IDM
100
100
100
100
100
100
1SP + EP-IDM
100
100
100
100
100
100
MSP + SIG
0
40
100
100
100
100
DINOv2
MSP
0
0
0
0
0
0
Appendix
Table 10 : CartPole success rates (%) with action contexts one and five. PR denotes a learned projector on top of the frozen encoder.
LQR (SR)
CEM (SR)
GBP (SR)
Encoder
Objective
Context 1
Context 5
Context 1
Context 5
Context 1
Context 5
From scratch
1SP + IDM
70
80
100
90
100
80
1SP + EP-IDM
80
80
100
100
80
80
1SP + IDM + MS-IDM
80
70
90
100
90
90
MSP + SIG
60
80
90
70
100
60
MSP + MS-IDM
80
80
100
100
90
90
Appendix
Table 11 : PointMaze success rates (%) with action contexts one and five.
Predictor
Objective
LQR (SR)
LQR (MFS)
CEM (SR)
CEM (held)
GBP (SR/held)
MLP
MSP + SIG
20%
0.571
20%
20%
0% / 0%
1SP + EP-IDM
0%
0.204
100%
100%
0% / 0%
Transformer
MSP + SIG
20%
0.461
0%
0%
0% / 0%
1SP + EP-IDM
80%
0.869
100%
100%
100% / 100%
Appendix
Table 12 : CartPole ablation: Action context one, only proprioceptive information (state).
Objective
Proprioception
LQR (SR)
CEM (SR)
CEM (held)
GBP (SR)
GBP (held)
MSP
Yes
0%
100%
100%
60%
40%
No
0%
0%
0%
0%
0%
MSP + PR-SIG
Yes
40%
90%
90%
60%
60%
No
0%
0%
0%
0%
0%
MSP + PR-IDM
Yes
100%
80%
80%
70%
70%
No
0%
0%
0%
0%
0%
Appendix
Table 13 : CartPole ablation with a frozen iBOT encoder and action context five, with and without proprioceptive inputs. PR denotes the learned projector. SR: success rate, held: fraction of trials in which stabilization is preserved.
Encoder
∥x0∥=0.005
∥x0∥=0.01
∥x0∥=0.02
∥x0∥=0.05
∥x0∥=0.10
∥x0∥=0.20
iBOT
1.07 ×
1.11 ×
1.07 ×
1.07 ×
1.02 ×
1.03 ×
DINOv2
25.45 ×
26.07 ×
26.38 ×
25.67 ×
21.83 ×
25.12 ×
Appendix
Table 14 : Ablation on the encoder gain ∥Δzt∥/∥Δxt∥ as a function of initial-state perturbation magnitude ∥x0∥ . We start from a perturbed upright equilibrium, the cart pole is stepped with zero action for 60 steps. We measure the mean latent displacement ∥Δzt∥=∥zt−zt−1∥ divided by the mean physical-state displacement ∥Δxt∥=∥xt−xt−1∥ .
Encoder
Objective
∥P(zeq,0)−zeq∥
DINOv2
1SP
5.569
MSP
7.215
1SP + PR-EP-IDM
0.4026
MSP + PR-SIG
2.515
MSP + PR-EP-IDM
0.343
iBOT
1SP
0.059
Appendix
Table 15 : Fixed-point residual of the learned predictor at the encoded upright CartPole equilibrium under zero action. The residual measures the one-step drift from this encoded equilibrium. PR denotes a learned projector on top of the frozen encoder.
Figure 23 : Synthetic open-loop stable LDS. This example considers a two-dimensional open-loop stable system. Panels descriptions are the same as those in Fig. 2 .
Figure 24 : Sign of ΔV=V(x′)−V(x) after one closed-loop step under the learned latent LQR controller, where V(x)=x⊤Sx and S≻0 is the ground-truth Lyapunov function matrix. We also depict the ΔV=0 contour (separatrix), i.e., the boundary separating states where the learned controller decreases V (blue) from those where it does not (red).
School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore · The University of Sheffield, Sheffield, United Kingdom · School of Artificial Intelligence (School of Software), Yanshan University, Qinhuangdao, China