Multimodal world models are often evaluated by whether different sensors produce similar representations. However, similar representations do not necessarily imply that the models make the same physical predictions, or that those representations can be reused when actions are combined in a new order. We study both questions through the physical responses predicted by a model. We first use the Cluster Haptic dataset to ask whether audio and acceleration can independently recover the behavior of the same surface from different observations. Predictions from the two sensors are substantially closer for the same surface than for different surfaces, with a 4.5× gap on average, while both also outperform an average-surface prediction. We then show that this agreement alone does not determine how familiar actions should compose. In a controlled elastoplastic system, shared step dynamics fit observed programs less accurately than a whole-program predictor but generalize better to unseen action orders, with the ranking reversing on both held-out transitions across three independent initializations. Fusing free-decay and hysteresis observations further improves prediction, with diagonal Gaussian beliefs yielding the lowest errors. Together, these results distinguish cross-sensor consistency, multimodal fusion, and generalization to new action orders as separate questions in evaluating multimodal physical representations.
Figures & tables
(a) Response prediction
Model
Joint ES ↓
NMSE ↓
Population prior
0.10239
1.0000
Free decay
0.07909
0.5492
Hysteresis
0.06594
0.4066
Rank-2 fused
0.06147
0.3193
Diagonal fused
0.05456
0.2873
Table 1: Fusion and belief parameterization for test AB/BA responses. Results on 64 test entities with the frozen executor used for fusion. NMSE is relative to the executed population prior; order-effect error is relative RMS error. Lower is better. Single-modality order-effect errors were not reported. Diagnostic controls are reported in Appendix C.5 .
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Development
Test
Modality
P0
P1
P0
P1
Audio
0.8636
0.8050
0.8602
0.8298
Acceleration
0.7328
0.7386
0.6873
0.6870
Population
1.0000
1.0000
1.0000
1.0000
Appendix
Table 2: Response prediction on Cluster Haptic. Repeat-1 NMMSE on 19 development and 19 test surfaces, each evaluated on 67 bridge queries. The population chart-center baseline equals one. Test targets enter no fitting or model selection.
Model
Seed
Selected update
Fitting MSE
Dev. MSE
Test MSE
Query-only
1001
15,000
0.038488
0.202754
0.916450
Shared-step
0
30,000
0.004363
0.000780
0.007462
Whole-program
0
15,000
0.000147
0.273157
1.086167
Shared-step
1
30,000
0.005254
0.001508
0.005883
Whole-program
1
30,000
0.000102
0.230814
0.724599
Shared-step
2
30,000
0.001918
0.001610
0.004869
Appendix
Table 3: Executor comparison before baseline normalization. MSE averages over entities, programs, 60 post-action steps, and three channels standardized using training responses. Selection uses development MSE. The fixed query-only row supplies a common denominator within each column.
Diagnostic
Observed value
Interpretation
Free-decay variation over Fy
0
Free decay is invariant to yield force.
Hysteresis variation over c
0
Settled hysteresis is invariant to damping.
Target-response RMS from changing k,c,Fy
k : 0.09509 c : 0.03925 Fy : 0.09041
All three physical parameters affect the test response.
Chart relative residual
0.102265
Three coordinates retain 98.95% of centered fitting-response variation.
Oracle NMSE, dev / test
0.0756 / 0.1844
Exact coordinates give lower error than the zero-coordinate prediction.
AB/BA order-effect RMS
0.62016
Changing action order changes the zero-input response.
Appendix
Table 4: Physical and representation diagnostics of the controlled system. Blind-direction values are maximum evidence variations over the indicated parameter, with the other parameters fixed. Factor-effect RMS compares adjacent parameter levels in normalized test responses. Learned-model diagnostics use the 30,000-update training budget, with the executor selected by development error and fusion from the rank-two-factor model. The information eigenvalue excludes prior precision; trace reduction is relative to the smaller mean posterior trace of the two single modalities.
World models, systems that generate what happens next given current environmental conditions, are increasingly being implemented with multi-modal generation in mind. However, generating multiple modalities simultaneously, such as visual simulations alongside physical state predictions in the form of text, introduces the risk of cross-modal inconsistency. Tested separately, both outputs may look convincing while still disagreeing: a model can calculate that a ball should rebound in one modality, then generate no rebound in another modality, to say nothing of diverging from real-world dynamics entirely. In this work we focus on two failures explicitly: \emph{Internal misalignment}, the disagreement between the world model's generated video and the same world model's prediction in a different modalities, and \emph{external misalignment} the disagreement between the world model's generation and an analytic physical environment. We derive common contracts of event, magnitude, timing, and construct a physics grounded pipeline to make comparisons measurable in both external and internal settings. We then ask whether progressively supplying the model's own contract (the A ladder for the internal setting) or a corrected physical contract (the B ladder for the external setting) closes the respective gaps. Across four mechanisms and 20 settings, we find that while language answers all 22 text probes correctly with respect to the true environment, the neutral video is often in disagreement, suggesting that the current unified backbones may not be capable of correct reasoning, internal consistency, and external physical fidelity all at once.
Geigh Zollicoffer, Minh Vu, Rajiv Ranasinghe +1
1Theoretical Division, Los Alamos National Laboratory · 2Computing and Artificial Intelligence Division, Los Alamos National Laboratory · 3Earth and Environmental Sciences Division, Los Alamos National Laboratory
Recent advances in Human Activity Recognition (HAR) from wearable sensors have shown that multi-modal deep learning models consistently outperform their uni-modal counterparts. Modalities can include IMUs, RGB cameras, audio signals, and others. One important aspect of multi-modal deep learning is the sensor fusion approach we apply. Over recent years, multiple fusion paradigms have been proposed for multi-modal HAR. However, to the best of our knowledge, no head-to-head comparison of these paradigms exists on a common multi-modal HAR benchmark dataset. To address this research gap, we systematically compare seven state-of-the-art sensor fusion methods on the recently released HARMES dataset, which comprises 61 hours of fully labeled IMU, audio, and ambient humidity data. The chosen dataset focuses on 15 household and personal hygiene activities of daily living (ADLs). By applying the seven different fusion techniques to a state-of-the-art multi-modal model architecture, we show that Gated Multi-modal Fusion achieves the highest macro F1-score (0.82), surpassing the concatenation-based late fusion HARMES paper baseline of 0.76 by +6pp under leave-one-participant-out evaluation. All code used in our experiments is made publicly available on GitHub.
Ahmed Mohamady, Robin Burchard, Kristof Van Laerhoven
Multimodal systems often benefit from combining information across language, sound, and visual streams, but this benefit is not guaranteed. A modality that is useful for one input may become distracting for another, and local feature responses within the same modality can disagree with evidence from other sources. This work investigates how to adjust multimodal representations before they are merged by a downstream predictor. We develop a compact calibration module that compares each modality with the others at the summary level, extracts cues of cross-source support and conflict, and converts these cues into instance-wise and dimension-wise modulation signals. The calibration is applied to the original modality features rather than to already fused representations, enabling the model to suppress misleading components, preserve weak but useful evidence, and emphasize responses that are better supported by the current multimodal context. The module is designed as a plug-in component and can be attached to different fusion backbones without changing their prediction heads. Across five benchmarks covering sentiment understanding, action recognition, audio-visual event detection, and audio-visual emotion classification, the proposed pre-combination calibration strategy improves performance under both sequence-based and convolutional fusion settings. Additional analyses under modality removal, synthetic corruption, training dynamics, and feature-level visualization show that calibrating signals before fusion can reduce interference from unreliable modalities and produce more stable multimodal optimization.
Jiyuan Liu, Liangwei Nathan Zheng, Wei Emma Zhang +2
Adelaide University Adelaide, Australia · Shandong University Jinan, China