Directed Temporal Representations for Offline Visual Control
Organizations: University of Sheffield · Shanghai Jiao Tong University · Shanghai University of Finance and Economics · University College London
Abstract
Predictive world models provide compact visual representations for control. Control requires a latent geometry aligned with temporal reachability rather than predictive similarity alone. We introduce Directed Temporal Representations for Control (DTRC), which learns such a geometry from offline visual trajectories on top of frozen LeWorldModel (LeWM) features. DTRC constructs a directed temporal quasimetric over the learned control representation. Short-range temporal offsets calibrate the distance scale. Bootstrapped targets extend temporal reachability across longer horizons. Action-conditioned consistency aligns the representation with local transition dynamics. The resulting distance estimates temporal reaching cost, and its change across a transition defines goal-relative temporal progress. We use this progress signal as a temporal critic for direct goal-conditioned policy learning. Model-assisted targets provide an additional training-time refinement under behavior-support and dynamics-agreement constraints. Across ten visual control tasks, DTRC achieves strong goal-conditioned control performance relative to planning and direct-policy baselines. Held-out diagnostics on the four LeWM tasks show consistent short-range temporal calibration, task-dependent long-range and directional structure, and positive transition-level progress. Temporal supervision improves the same flow-policy parameterization across all four LeWM tasks, while the resulting policy acts directly without iterative trajectory search at test time.
Figures & tables
| Method | PointMaze | Cube-Double | Scene | AntMaze-Medium | HumanoidMaze-Medium | Puzzle-3x3 |
|---|---|---|---|---|---|---|
| LeWM | ||||||
| GCBC | ||||||
| GC-IQL | ||||||
| HILP | ||||||
| GC-IDM | ||||||
| DTRC (ours) |
| Task | MAE 1:8 | Reverse Forward | Positive progress | |
|---|---|---|---|---|
| (steps) | (%) | (%) | ||
| PushT | 1.55 | 0.845 | 92.76 | 81.41 |
| TwoRoom | 1.51 | 0.800 | 57.65 | 65.78 |
| Reacher | 1.62 | 0.510 | 50.28 | 59.41 |
| Cube | 0.74 | 0.221 | 52.46 | 72.35 |
| Configuration | PushT | TwoRoom | Reacher | Cube |
|---|---|---|---|---|
| Flow BC | ||||
| + temporal supervision |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Network | Input | Hidden widths | Output |
|---|---|---|---|
| Representation | |||
| Symmetric head | |||
| Order-sensitive head | |||
| Transition head | |||
| Dynamics member | |||
| Policy velocity |
| Hyperparameter | Value |
| Optimization | |
| Training iterations | |
| Minibatch size | |
| Optimizer | Adam |
| Learning rate | (constant) |
| Adam moment coefficients | |
| Configuration | PushT | TwoRoom | Reacher | Cube |
|---|---|---|---|---|
| Flow BC | ||||
| Flow BC + temporal supervision | ||||
| DTRC w/o behavior support | ||||
| DTRC w/o directional residual | ||||
| DTRC (full) |
| Goal offset | |||
|---|---|---|---|
| Method | 25 | 50 | 100 |
| LeWM | |||
| GCBC | |||
| GC-IQL | |||
| HILP | |||
| GC-IDM | |||