Existing multimodal scaling laws fit multimodality terms empirically after testing and never vary how much data is multimodally paired at fixed data budgets. We investigate how, under the same total data per modality, changing the number of paired data affects loss curves in multimodal classification tasks. We train models in three different environments and run experiment sweeps varying data sizes and pairing budget. Pairing ratios have a dramatic impact on loss and this impact is directly tied to how much information synergy the task contains. Only paired data is able to reduce synergistic loss, while unpaired data can reduce redundant or unimodal information up until unimodal floors. Unlike traditional scaling laws where loss drops immediately in power law decay, synergy acquisition is gated, requiring a critical threshold of paired data before synergistic loss falls at all. We introduce a new family of multimodal scaling laws where total data-attributable loss is the sum of four individual power laws corresponding to the four different information channels of redundancy, a unique channel per modality, and synergy, and show how this law is both more theoretically sound and empirically valid across our experiments. This law predicts multimodal loss in our experiments more accurately than existing laws, with 3.2% error on fit tests versus 10.4% error for the best pairing extension of published laws.
Figures & tables
Figure 1: Illustration of each environment along with an example and tested architecture.
Digits
VQA-v2
NLVR2
image
audio
image
question
photos
sentence
Experimental details
architecture
CNN + 4-way head ( 476 k)
Frozen CLIP + head ( 5.3 M)
Finetuned LXMERT ( 210 M)
chance loss L0
ln10≈2.303
ln3129≈8.048
ln2≈0.693
D
8 k, 16 k, 32 k, 64 k, 128 k, 256 k
0 , 4 k … 256 k
4 k, 8 k, 16 k, 32 k, 64 k
C ( k(D1+D2) steps)
k∈{10,20,40}
k∈Z+≤80
k∈{1,2,3}
Table 1: Experimental details, unimodal data curve variables, and PID terms for each environment. Compute C represents number of training steps multiplied by batch size.
Law
k
Γ
∂Γ/∂D=0
L(N,D1,D2,np)
at p=1
Hoffmann Unimodal
4
≤0
✓
EN+Bg(D1+D2)
EN+Bg(2D)
Hoffmann on unique D
4
≤0
✓
EN+Bg(Dtot)
EN+Bg(D)
Hoffmann on D1 and D2
7
=0
×
EN+B1g1(D1)+B2g2(D2)
EN+B1g1(D)+B2g2(D)
Aghajanyan +Cp
8
=0
×
EN+B1g1(D1)+B2g2(D2)+Cp
EN+B1g1(D)+B2g2(D)+C
Shukor mixture
6
<0
✓
EN+(C1p+C2(1−p))−1+Bg(Dtot)
EN+C1−1+Bg(D)
Ye mixture
7
any
✓
EN+Cexp(α2n2/Dtot+αpnp/Dtot)+Bg(Dtot)
EN+Ceαp+Bg(D)
Table 2: Extensions of existing scaling laws (Top) and the introduced four-channel law and variants (Bottom). Existing laws reduce to their original equations at p=1 . Each of the variations of the four-channel law can be applied independently, resulting in 16 four-channel law combinations. k is the number of fitted parameters for the corresponding law, Dtot=D1+D2−np , ni=Di−np , and Wa are measured PID weights. Γ=−∂2L/∂D1∂D2 is taken at fixed p ; for the four-channel family its sign is that of the base law unless a switch changes it. Γ and ∂Γ/∂D are discussed in § 5.3 .
Digits
VQA
NLVR2
Overall
Law
k
D↑
p=1
R2
D↑
p=1
R2
D↑
p=1
R2
D↑
p=1
R2
Hoffmann Unimodal
4
13.9%
30.4%
0.00
8.2%
8.1%
0.00
1.8%
6.6%
0.00
8.0%
15.0%
0.00
Hoffmann on unique D
4
14.9%
37.1%
-0.12
10.5%
13.1%
-1.26
1.7%
6.9%
-0.45
9.0%
19.0%
-0.61
Hoffmann on D1 and D2
7
11.5%
27.3%
0.00
2.7%
6.7%
0.00
1.8%
6.6%
0.00
5.3%
13.5%
0.00
Aghajanyan +Cp
8
11.5%
20.4%
0.39
2.3%
5.0%
0.40
1.5%
6.0%
0.22
5.1%
10.4%
0.34
Shukor mixture
6
9.6%
29.1%
0.46
8.2%
9.4%
0.02
1.4%
6.2%
-0.09
6.4%
14.9%
0.13
Table 3: Paired scaling law scores on the three fit tests: percent error for D↑ and p=1 (lower is better), and within-cell R2 (higher is better).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Digits
VQA
NLVR2
Overall
Law
k
LOO
Full
LOO
Full
LOO
Full
LOO
Full
Hoffmann Unimodal
4
5.9%
5.6%
3.9%
3.7%
1.0%
1.3%
3.6%
3.5%
Hoffmann on unique D
4
5.2%
4.5%
4.0%
3.8%
1.0%
1.2%
3.4%
3.1%
Hoffmann on D1 and D2
7
4.2%
3.6%
1.1%
1.0%
1.2%
1.3%
2.2%
2.0%
Aghajanyan +Cp
8
5.7%
4.6%
1.2%
1.0%
0.9%
0.8%
2.6%
2.2%
Shukor mixture
6
4.1%
3.9%
4.6%
4.3%
0.8%
1.1%
3.2%
3.1%
Appendix
Table 4: Paired scaling law scores of the two interpolation tests: Leave-One-Out and the Full test.