Existing multimodal scaling laws fit multimodality terms empirically after testing and never vary how much data is multimodally paired at fixed data budgets. We investigate how, under the same total data per modality, changing the number of paired data affects loss curves in multimodal classification tasks. We train models in three different environments and run experiment sweeps varying data sizes and pairing budget. Pairing ratios have a dramatic impact on loss and this impact is directly tied to how much information synergy the task contains. Only paired data is able to reduce synergistic loss, while unpaired data can reduce redundant or unimodal information up until unimodal floors. Unlike traditional scaling laws where loss drops immediately in power law decay, synergy acquisition is gated, requiring a critical threshold of paired data before synergistic loss falls at all. We introduce a new family of multimodal scaling laws where total data-attributable loss is the sum of four individual power laws corresponding to the four different information channels of redundancy, a unique channel per modality, and synergy, and show how this law is both more theoretically sound and empirically valid across our experiments. This law predicts multimodal loss in our experiments more accurately than existing laws, with 3.2% error on fit tests versus 10.4% error for the best pairing extension of published laws.
Figures & tables
Figure 1: Illustration of each environment along with an example and tested architecture.
Digits
VQA-v2
NLVR2
image
audio
image
question
photos
sentence
Experimental details
architecture
CNN + 4-way head ( 476 k)
Frozen CLIP + head ( 5.3 M)
Finetuned LXMERT ( 210 M)
chance loss L0
ln10≈2.303
ln3129≈8.048
ln2≈0.693
D
8 k, 16 k, 32 k, 64 k, 128 k, 256 k
0 , 4 k … 256 k
4 k, 8 k, 16 k, 32 k, 64 k
C ( k(D1+D2) steps)
k∈{10,20,40}
k∈Z+≤80
k∈{1,2,3}
Table 1: Experimental details, unimodal data curve variables, and PID terms for each environment. Compute C represents number of training steps multiplied by batch size.
Law
k
Γ
∂Γ/∂D=0
L(N,D1,D2,np)
at p=1
Hoffmann Unimodal
4
≤0
✓
EN+Bg(D1+D2)
EN+Bg(2D)
Hoffmann on unique D
4
≤0
✓
EN+Bg(Dtot)
EN+Bg(D)
Hoffmann on D1 and D2
7
=0
×
EN+B1g1(D1)+B2g2(D2)
EN+B1g1(D)+B2g2(D)
Aghajanyan +Cp
8
=0
×
EN+B1g1(D1)+B2g2(D2)+Cp
EN+B1g1(D)+B2g2(D)+C
Shukor mixture
6
<0
✓
EN+(C1p+C2(1−p))−1+Bg(Dtot)
EN+C1−1+Bg(D)
Ye mixture
7
any
✓
EN+Cexp(α2n2/Dtot+αpnp/Dtot)+Bg(Dtot)
EN+Ceαp+Bg(D)
Table 2: Extensions of existing scaling laws (Top) and the introduced four-channel law and variants (Bottom). Existing laws reduce to their original equations at p=1 . Each of the variations of the four-channel law can be applied independently, resulting in 16 four-channel law combinations. k is the number of fitted parameters for the corresponding law, Dtot=D1+D2−np , ni=Di−np , and Wa are measured PID weights. Γ=−∂2L/∂D1∂D2 is taken at fixed p ; for the four-channel family its sign is that of the base law unless a switch changes it. Γ and ∂Γ/∂D are discussed in § 5.3 .
Digits
VQA
NLVR2
Overall
Law
k
D↑
p=1
R2
D↑
p=1
R2
D↑
p=1
R2
D↑
p=1
R2
Hoffmann Unimodal
4
13.9%
30.4%
0.00
8.2%
8.1%
0.00
1.8%
6.6%
0.00
8.0%
15.0%
0.00
Hoffmann on unique D
4
14.9%
37.1%
-0.12
10.5%
13.1%
-1.26
1.7%
6.9%
-0.45
9.0%
19.0%
-0.61
Hoffmann on D1 and D2
7
11.5%
27.3%
0.00
2.7%
6.7%
0.00
1.8%
6.6%
0.00
5.3%
13.5%
0.00
Aghajanyan +Cp
8
11.5%
20.4%
0.39
2.3%
5.0%
0.40
1.5%
6.0%
0.22
5.1%
10.4%
0.34
Shukor mixture
6
9.6%
29.1%
0.46
8.2%
9.4%
0.02
1.4%
6.2%
-0.09
6.4%
14.9%
0.13
Table 3: Paired scaling law scores on the three fit tests: percent error for D↑ and p=1 (lower is better), and within-cell R2 (higher is better).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Digits
VQA
NLVR2
Overall
Law
k
LOO
Full
LOO
Full
LOO
Full
LOO
Full
Hoffmann Unimodal
4
5.9%
5.6%
3.9%
3.7%
1.0%
1.3%
3.6%
3.5%
Hoffmann on unique D
4
5.2%
4.5%
4.0%
3.8%
1.0%
1.2%
3.4%
3.1%
Hoffmann on D1 and D2
7
4.2%
3.6%
1.1%
1.0%
1.2%
1.3%
2.2%
2.0%
Aghajanyan +Cp
8
5.7%
4.6%
1.2%
1.0%
0.9%
0.8%
2.6%
2.2%
Shukor mixture
6
4.1%
3.9%
4.6%
4.3%
0.8%
1.1%
3.2%
3.1%
Appendix
Table 4: Paired scaling law scores of the two interpolation tests: Leave-One-Out and the Full test.
A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone. While most approaches operate at the architectural level through larger or more complex fusion models, we propose a complementary axis: shaping the training objective itself. Standard training often emphasizes unimodal or redundant information, falling short on examples that require cross-modal reasoning. We formalize multimodal synergy through information theory and introduce the Synergistic Information Bottleneck (SynIB), a scalable objective that targets synergy directly. To prioritize learning synergy, SynIB motivates the model to predict accurately from all modalities while penalizing confidence when information from any modality is withheld. Alongside the standard task loss, the model runs forward passes with one modality masked at a time and is penalized for remaining confident, which would indicate reliance on unimodal cues rather than cross-modal interactions. We validate SynIB in two regimes. On synthetic XOR tasks where the ground-truth synergy is known by construction, standard training fails to recover it while SynIB does. On five real-world benchmarks, including three MultiBench affective tasks, Hateful Memes with CLIP-ViT and DeBERTa backbones, and a controllable irony extension of CREMA-D we introduce, SynIB improves accuracy on synergy-dependent examples by up to 7.8% and overall accuracy by up to 3.8%.
Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen +4
Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities. Most current approaches are built on the assumption of a deterministic one-to-one alignment between modalities. However, this oversimplifies real-world multimodal relationships, where their nature is inherently many-to-many. The many-to-many property, or multiplicity, is not a side-effect of noise or annotation error, but an inevitable outcome of intra-modal variability, representational asymmetry, and task-dependent ambiguity in multimodal tasks. We argue that multiplicity is a fundamental bottleneck that affects all stages of the multimodal learning pipeline: from data construction to model training and evaluation benchmarks. By formalizing its causes and consequences, we demonstrate how ignoring multiplicity leads to training uncertainty, unreliable evaluation, and degraded dataset quality. This position paper calls for new research directions on multimodal learning, including multiplicity-aware learning frameworks and dataset construction and evaluation protocols.
Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds and when each fails --- a gap that leaves practitioners, especially in scientific domains with heterogeneous instruments and multiple levels of measurement, unable to diagnose why standard methods underperform the best single modality. We study both objectives under a spiked signal-plus-noise model with structured cross-modal nuisance correlation, the ingredient that breaks the classical recovery guarantees, and derive separation ratios that expose complementary failure modes: alignment whitens each modality and fails when nuisance is strongly correlated across views; prediction encodes whatever is cross-predictable through a one-sided whitening, with recovery governed by source-modality quality. The resulting phase diagram partitions multimodal problems into four regimes --- Both, CA only, CP only, and Neither --- refined by a recovery count that separates partial recovery from complete failure. We present a data-driven procedure to locate real-world datasets in this diagram using a small labeled subsample, identifying the preferred objective and prediction direction before any cross-modal training, and identifying when no objective in the CA/CP family can improve on the stronger modality alone. Experiments on synthetic data, stereo-vision benchmarks, image--caption pairs, and two real scientific domains --- astronomy and single-cell multi-omics --- validate the predictions in the nonlinear regime, including both faces of the Neither regime. Code to reproduce the results is available at https://github.com/IlayMalinyak/mm_align_vs_pred.