cs.LGSep 28, 2026

Paired Multimodal Scaling Laws

Authors: Marcus Ma, Shrikanth Narayanan

Organizations: University of Southern California

Abstract

Existing multimodal scaling laws fit multimodality terms empirically after testing and never vary how much data is multimodally paired at fixed data budgets. We investigate how, under the same total data per modality, changing the number of paired data affects loss curves in multimodal classification tasks. We train models in three different environments and run experiment sweeps varying data sizes and pairing budget. Pairing ratios have a dramatic impact on loss and this impact is directly tied to how much information synergy the task contains. Only paired data is able to reduce synergistic loss, while unpaired data can reduce redundant or unimodal information up until unimodal floors. Unlike traditional scaling laws where loss drops immediately in power law decay, synergy acquisition is gated, requiring a critical threshold of paired data before synergistic loss falls at all. We introduce a new family of multimodal scaling laws where total data-attributable loss is the sum of four individual power laws corresponding to the four different information channels of redundancy, a unique channel per modality, and synergy, and show how this law is both more theoretically sound and empirically valid across our experiments. This law predicts multimodal loss in our experiments more accurately than existing laws, with 3.2% error on fit tests versus 10.4% error for the best pairing extension of published laws.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning

    May 12, 2026Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen +4Multimodal LearningCross-Modal

  2. When to Align, When to Predict: A Phase Diagram for Multimodal Learning

    Jun 9, 2026Ilay Kamai, Hugues Van Assel, Aviv Regev +2Multimodal LearningMultimodal Representations