cs.LGAug 19, 2026

Beyond Multimodal Alignment: Shared Physical Representations Across Sensors and Action Orders

Authors: Kaizhen Tan, Xin Xu, Siru Tao, Yixiao Li, Hanzhe Hong, Yang Feng, Heqing Du

Organizations: New York University, New York, NY, USA · Carnegie Mellon University, Pittsburgh, PA, USA · Columbia University, New York, NY, USA

Abstract

Multimodal world models are often evaluated by whether different sensors produce similar representations. However, similar representations do not necessarily imply that the models make the same physical predictions, or that those representations can be reused when actions are combined in a new order. We study both questions through the physical responses predicted by a model. We first use the Cluster Haptic dataset to ask whether audio and acceleration can independently recover the behavior of the same surface from different observations. Predictions from the two sensors are substantially closer for the same surface than for different surfaces, with a 4.5×4.5\times gap on average, while both also outperform an average-surface prediction. We then show that this agreement alone does not determine how familiar actions should compose. In a controlled elastoplastic system, shared step dynamics fit observed programs less accurately than a whole-program predictor but generalize better to unseen action orders, with the ranking reversing on both held-out transitions across three independent initializations. Fusing free-decay and hysteresis observations further improves prediction, with diagonal Gaussian beliefs yielding the lowest errors. Together, these results distinguish cross-sensor consistency, multimodal fusion, and generalization to new action orders as separate questions in evaluating multimodal physical representations.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling

    Sep 13, 2026Geigh Zollicoffer, Minh Vu, Rajiv Ranasinghe +1Cross-Modal ConflictsMultimodal Model

  2. A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset

    Jun 26, 2026Ahmed Mohamady, Robin Burchard, Kristof Van LaerhovenMultimodal FusionHuman Activity Recognition

  3. Before Fusion, Ask What to Keep: Contextual Calibration of Multimodal Signals

    Jun 1, 2026Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang +2Multimodal FusionMultimodal Representations