Mutual information (MI) provides an objective for suppressing or encouraging statistical dependence in implicit generative models. However, direct MI evaluation is challenging in implicit models due to typically intractable densities. A remedy is estimating the generator gradient from the difference between conditional and marginal scores. This score difference can, in turn, be estimated by differentiating a log density ratio learned through classification. This construction nevertheless faces two difficulties: (i) singular distributions need not admit the required score functions, and (ii) poor overlap can hinder density-ratio estimation. We therefore introduce Spread Mutual Information (SMI), a weighted integral of MI across noise levels obtained by applying a common spreading kernel to the generated variable. Gaussian spreading yields smooth, strictly positive conditional and marginal densities, extending the gradient construction to distributions that may originally be singular. Across a variaty of experiments, SMI consistently achieves effective dependence control among MI-based methods and remains competitive with established task-specific approaches.
Figures & tables
Figure 1: Gaussian spreading near a singular limit. Left pair: Joint and product-of-marginals samples at b=0 , before (leftmost) and after (middle) spreading with σ=0.5 . Right: Relative variance of the pathwise estimator of ∂aI(X;Y) using exact scores.
Figure 2: Synthetic experiments. (a) Conditional supports x2=±a and optimisation of their separation 2a . (b,c) MI minimisation and maximisation results, pooled over Gaussian, cubic, asinh, signed-power and spiral data; d∈{20,100,200} ; a shared correlation coefficient or one per latent coordinate pair; and ten seeds.
Estimator
Score difference
Parameter gradient
SMI
0.408±0.030
0.275±0.031
DSM scores
2.608±0.376
6.225±1.834
Table 1: Gradient accuracy. Relative errors ( ↓ ), mean ± SD over ten seeds.
(a) 30% MCAR
(b) 60% MCAR
(c) 80% MCAR
Method
MMD 2 rank
RMSE rank
MMD 2 rank
RMSE rank
MMD 2 rank
RMSE rank
MI-based methods
SMI (Ours)
5.43±0.31
7.63±0.15
4.10±0.10
7.63±0.06
2.87±0.21
7.33±0.35
CLUB
12.30±0.66
9.93±0.21
11.97±0.21
7.37±0.12
11.00±0.20
5.33±0.12
NWJ
11.90±0.20
13.63±0.23
8.43±0.40
11.80±0.20
6.47±0.83
10.53±0.95
MINE
10.33±0.83
12.90±0.40
7.43±0.32
11.27±0.15
7.40±0.66
10.40±0.53
Table 2: Missing data imputation. Ranks across all methods ( ↓ ), averaged over ten datasets at each missingness rate; mean ± SD over three seeds. Failed runs rank last.
Figure 3: Waterbirds. Training examples from the four bird–background groups.
Worst-group
Leakage (points, ↓ )
Method
accuracy (%, ↑ )
Linear
MLP
MI-based methods
SMI (Ours)
71.68±1.33
18.19±5.35
9.20±5.49
MINE
70.87±1.16
26.21±2.42
22.23±4.00
SMILE
71.43±1.83
25.34±3.15
17.69±1.76
InfoNCE
71.43±1.40
25.59±3.14
18.12±4.09
Table 3: Waterbirds. Test results; mean ± SD over five training seeds.
Figure 4: Edges-to-photos translation. Input, paired photo and two SMI outputs (different latents).
Latent-code prediction R2 ( ↑ )
Method
Unperturbed
JPEG (75)
Resize
MI-based methods
SMI (Ours)
0.798±0.008
0.623±0.045
0.726±0.032
InfoNCE
0.757±0.004
0.588±0.032
0.623±0.071
NWJ
0.788±0.008
0.596±0.047
0.682±0.033
FLO
0.725±0.001
0.503±0.028
0.619±0.025
Table 4: Edges2Shoes. Latent-code prediction R2 ( ↑ ); mean ± SD over three training seeds.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Minimisation ↓
Maximisation (target: 50)
Method
d=20
d=100
d=200
d=20
d=100
d=200
SMI (Ours)
(1.36±0.366)×10−8
(7.83±11.6)×10−8
(1.00±2.74)×10−6
50.007±0.006
50.205±0.111
50.230±0.197
CLUB
(1.12±0.373)×10−5
(1.20±0.362)×10−4
0.973±0.025
50.006±0.003
50.268±0.057
50.095±0.058
NWJ
(2.44±1.56)×10−5
(6.59±4.85)×10−5
(4.84±3.08)×10−4
n.f.
n.f.
50.430±0.316
MINE
n.f.
(6.55±7.41)×10−5
(1.00±0.827)×10−4
n.f.
n.f.
n.f.
InfoNCE
(1.90±0.991)×10−5
(6.26±5.40)×10−5
(6.90±4.34)×10−5
21.391±0.246
49.526±0.556
50.009±0.007
Appendix
Table 5: Gaussian MI optimisation. Shared correlation ( r=d ); mean ± SD in nats over ten runs. n.f.: at least one numerical failure.
Figure 5: Gaussian data. Outcomes by dimension, pooled over one shared correlation parameter or one per coordinate pair (ten runs each).
Figure 6: Cubic data. Outcomes by dimension, pooled over one shared correlation parameter or one per coordinate pair (ten runs each).
Figure 7: Asinh data. Outcomes by dimension, pooled over one shared correlation parameter or one per coordinate pair (ten runs each).
Figure 8: Signed-power data. Outcomes by dimension, pooled over one shared correlation parameter or one per coordinate pair (ten runs each).
Figure 9: Spiral data. Outcomes by dimension, pooled over one shared correlation parameter or one per coordinate pair (ten runs each).
Figure 10: Information across noise levels. Exact-density evaluations. (a) 2τE[vτ2] (lines) and −∂I/∂logτ (markers), with τ=σ2 . (b) Multiscale SMI for σ∼LogUniform(0.01η,η) .
Dataset
Source
Rows
Features
Licence
Blood
https://www.openml.org/d/46913
748
4
CC BY 4.0
Concrete
https://www.openml.org/d/46917
1,030
8
CC BY 4.0
Fish
https://www.openml.org/d/46954
907
6
CC BY 4.0
Hazelnut
https://www.openml.org/d/46930
2,400
30
CC BY-SA
Houses
https://www.openml.org/d/46934
20,640
8
Public
Maternal
https://www.openml.org/d/46941
1,014
6
CC BY 4.0
Appendix
Table 6: TabArena datasets. All rows; numerical feature counts exclude targets. Licences follow the linked OpenML records.
Figure 11: Missing data imputation. MMD 2 and RMSE at 30%, 60% and 80% MCAR; individual runs and mean ± SD for groups with three successful fits.
Figure 12: Missing data imputation (continued).
Figure 13: Missing data imputation (continued).
Figure 14: Missing data imputation (continued).
Accuracy (%, ↑ )
Leakage (points, ↓ )
Method
Overall
Worst group
Linear
MLP
MI-based methods
SMI (Ours)
89.96±0.32
71.68±1.33
18.19±5.35
9.20±5.49
MINE
90.68±0.32
70.87±1.16
26.21±2.42
22.23±4.00
SMILE
90.62±0.36
71.43±1.83
25.34±3.15
17.69±1.76
InfoNCE
90.75±0.22
71.43±1.40
25.59±3.14
18.12±4.09
Appendix
Table 7: Waterbirds. Test accuracy and background leakage; mean ± SD over five training runs.
Method
KID ( ↓ )
Edge F1 ( ↑ )
ResNet-50 MLP R2 ( ↑ )
MI-based methods
SMI (Ours)
(1.587±0.145)×10−2
0.679±0.004
0.584±0.008
InfoNCE
(1.293±0.084)×10−2
0.685±0.003
0.523±0.007
NWJ
(1.481±0.064)×10−2
0.679±0.002
0.573±0.016
FLO
(1.177±0.014)×10−2
0.690±0.002
0.470±0.005
KNIFE
(1.451±0.154)×10−2
0.684±0.002
0.533±0.011
Appendix
Table 8: Edges2Shoes additional results. Reporting-set metrics at 20,000 generator updates; mean ± SD over three training runs.
Figure 15: Computational costs. Iteration time and peak allocated GPU memory across dimension, minibatch size and correlated coordinates; median and range over three repetitions.
Method
Objective or gradient
Time O(⋅)
Peak memory O(⋅)
SMI (Ours)
Spread log-ratio gradient
(Naux+1)B(dh+Nℓh2)
dh+Nℓh2+B(d+Nℓh)
CLUB
Sampled Gaussian contrast
(Naux+1)B(dh+Nℓh2)
dh+Nℓh2+B(d+Nℓh)
NWJ
Variational KL bound
(Naux+1)B2(dh+Nℓh2)
dh+Nℓh2+B2(d+Nℓh)
MINE
DV; moving normaliser
(Naux+1)B2(dh+Nℓh2)
dh+Nℓh2+B2(d+Nℓh)
InfoNCE
Contrastive bound
(Naux+1)B2(dh+Nℓh2)
dh+Nℓh2+B2(d+Nℓh)
SMILE
JS gradient; DV readout
(Naux+1)B2(dh+Nℓh2)
dh+Nℓh2+B2(d+Nℓh)
Appendix
Table 9: MI control methods and computational costs. Estimator costs per generator update for the evaluated synthetic implementations.