Physical systems are naturally parameterized in terms of geometric entities and their transformations. While exact equivariance can be a powerful inductive bias, real-world applications frequently exhibit approximate or broken symmetries in practice. We introduce object-centric world models that provide a soft geometric inductive bias without enforcing exact symmetries. By embedding object states as Clifford multivectors, our models encourage physically meaningful transformations while retaining the expressivity needed to handle asymmetric dynamics. We evaluate our approach on 2D rigid-body dynamics, 3D charged-particle systems, and real-world driving trajectories. Compared to both unstructured baselines and strictly equivariant models, our soft Clifford transformer achieves better long-horizon fidelity, particularly in regimes with broken symmetries. These results suggest that geometric algebra offers an effective middle ground, delivering sample-efficient dynamics without the inflexibility of hard mathematical constraints.
Figures & tables
Figure 1 : Rollout RMSE and average displacement error (ADE) for autoregressive object-centric world models trained on 2D rigid body physics with collisions (left), 3D charged particles in harmonic potential (middle), and real-world multi-agent driving scenarios from the Waymo Open Motion Dataset (WOMD) (right). Our Clifford transformer with a soft geometric inductive bias ( S-CliffordTransformer ) outperforms the vanilla transformer ( Transformer ) and the equivariant Clifford transformer ( E-CliffordTransformer ) in all data domains.
Figure 2
Blocks ( B )
Heads ( H )
MV ( C )
Np
10
8
24
1.5M
Table 1 : Clifford transformers hyperparameters with number of multivector channel C and number of trainable parameters Np .
Figure 4Figure 5Figure 6
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8 : Example rollout of ground truth (bottom) and model prediction (top) of frames 1 (left) to 11 (right) using our S-CliffordTransformer. The model predicts position, velocity, angle and angular velocity of the 10 rigid body polygons colliding with both the other dynamic bodies and the static environment with gravity.
Figure 8Figure 9
E-CliffordTransformer
S-CliffordTransformer
Transformer
Agent Type
SDC
0.0708
0.0334
0.1023
Non-SDC
4.6867
4.5321
4.8162
Appendix
Table 2 : Next-step XY position RMSE on WOMD (4096 scenarios) by agent type. Bold indicates best.
E-CliffordTransformer
S-CliffordTransformer
Transformer
Variable
x
0.0369
0.0163
0.0703
y
0.0544
0.0254
0.0722
xy
0.0708
0.0334
0.1023
vx
0.0314
0.0303
0.0664
vy
0.0651
0.0682
0.0775
Appendix
Table 3 : Per-variable RMSE breakdown for SDC agents on WOMD (4096 scenarios). Bold indicates best.
E-CliffordTransformer
S-CliffordTransformer
Transformer
Variable
x
3.2378
3.1310
3.4162
y
3.1245
3.0238
3.3811
xy
4.6867
4.5321
4.8162
vx
0.4371
0.4326
0.7522
vy
0.4106
0.4010
0.7788
Appendix
Table 4 : Per-variable RMSE breakdown for non-SDC agents on WOMD (4096 scenarios). Bold indicates best.
Figure 11 : Training and validation loss curves on the Waymo Open Motion Dataset for the E-CliffordTransformer (left), S-CliffordTransformer (middle) and Transformer (right) across dataset sizes (4096 and 16384 scenarios).
Figure 12 : Additional ablations on the 3D charged particle system. Left: varying context length (number of input frames). Right: varying prediction horizon (number of steps predicted). The S-CliffordTransformer maintains robust performance across all settings.
Figure 13 : Validation loss and 10 frame rollout RMSE during training for different number of transformer blocks (5, 10, 20) for our {S, S-Ad}-CliffordTransformer models and baseline Transformer . Trained on 100 episodes with sequence length of 16.
Figure 14 : Validation loss and 10 frame rollout RMSE during training for different number of transformer heads (4, 8, 16) for our {S, S-Ad}-CliffordTransformer models and baseline Transformer . Trained on 100 episodes with sequence length of 16.
Figure 15 : Validation loss and 10 frame rollout RMSE during training for different number of multivector channels (12, 24, 48) for our {S, S-Ad}-CliffordTransformer models. Trained on 100 episodes with sequence length of 16.
Figure 16 : Validation loss and 10 frame rollout RMSE during training for different embedding dimensions (32, 64, 128) for the baseline Transformer . Trained on 100 episodes with sequence length of 16.
Figure 17 : Example rollout using baseline transformer trained on 1000 episodes of 10 circles using sequence length 16. For each timestep, model prediction is shown in the top panel, and the ground truth environment simulation in the bottom panel.
Figure 18 : Example rollout using S-Ad-CliffordTransformer , using adjoint linear layers, trained on 1000 episodes of 10 circles using sequence length 16. For each timestep, model prediction is shown in the top panel, and the ground truth environment simulation in the bottom panel.
Figure 19 : Example rollout using baseline transformer trained on 1000 episodes of 4 circles using sequence length 16. For each timestep, model prediction is shown in the top panel, and the ground truth environment simulation in the bottom panel.
Figure 20 : Example rollout using S-Ad-CliffordTransformer , using adjoint linear layers, trained on 1000 episodes of 4 circles using sequence length 16. For each timestep, model prediction is shown in the top panel, and the ground truth environment simulation in the bottom panel.
Figure 21 : Example rollout using Equi-CliffordTransformer , trained on 1000 episodes of 10 rectangles using sequence length 2. For each timestep, model prediction is shown in the top panel, and the ground truth environment simulation in the bottom panel.
Figure 22 : Example rollout using S-CliffordTransformer , trained on 1000 episodes of 10 rectangles using sequence length 2. For each timestep, model prediction is shown in the top panel, and the ground truth environment simulation in the bottom panel.
Figure 23 : Example rollout using the baseline Transformer , trained on 1000 episodes of 10 rectangles using sequence length 2. For each timestep, model prediction is shown in the top panel, and the ground truth environment simulation in the bottom panel.