Standard information bottleneck (IB) regularization constrains representations via a single scalar I(Z;X), implicitlytreating all information as homogeneous. However, a single global compression control couples label-relevant structurewith residual within-condition variation, rather than regulating their allocation independently, allowing nuisanceinformation to persist in learned representations. For example, in medical imaging applications, residual variation oftenstems from acquisition conditions, background factors, or subject-specific appearance. This issue becomes particularlypronounced in data-limited settings, where models tend to overfit such variation, hindering generalization. While existingregularization methods can stabilize training, control capacity, or shape representation geometry, they do not explicitlyseparate nuisance-like variation from task-supporting structure. To address this limitation, we revisit IB from a structuredperspective based on a label-induced partition, where condition-level structure and within-condition information playdistinct roles. This leads to a dual-bottleneck formulation: a standard KL term controls global information capacity, while aconditional KL term targets within-condition information. We show that the conditional KL admits an exact decompositioninto a within-condition information term and a prior-mismatch term, explaining its alignment with the design objective.With a simplex-structured conditional prior, the method provides controllable latent geometry and integrates seamlesslyinto existing pipelines. Experiments on classification and segmentation show the clearest gains in low-data classificationand consistent improvements across dense prediction benchmarks.
Figures & tables
Method
Global
Cond.
Geometry / prior
Main distinction
VIB / StdKL
Jstd
–
global prior
only global bottleneck
CondKL (cond-only)
–
Jcond
conditional prior
no label-induced dual control
Center / prototype
–
implicit
empirical centers
feature-space geometry
SupCon
–
implicit
contrastive pairs
pairwise geometry
Generic regularizers
generic
–
none
orthogonal heuristics
DualKL (ours)
Jstd
Jcond
simplex conditional prior
joint two-axis control
Table 1: Objective-level comparison. Jstd and Jcond denote global and label-conditioned KL controls.
Figure 1: BloodMNIST feature heterogeneity. Points denote low-level statistics with sample separation on the x-axis and label separation on the y-axis. Lower-right points correspond to high variation but weak label relevance.
Figure 2: Information-plane view of representation control. Panel (a) illustrates the decomposition I(Z;X)=I(Z;C)+I(Z;X∣C) . Panels (b)–(d) schematically compare global-only, conditional-only, and dual control directions.
Figure 3: Visual illustration of complementary control in latent space. Global control mainly contracts overall spread, conditional control tightens within-condition structure, and dual control combines both effects in this illustrative example.
Figure 4: Engineering instantiation of the dual-bottleneck objective. The encoder produces a stochastic latent representation Z , the task head optimizes the supervised loss, and two KL surrogates implement global and condition-aware information control.
Figure 5: BloodMNIST, 1% training data. No Reg (left two panels) versus DualKL (right two panels), shown in PCA and latent-statistics views. Full four-way comparisons are in Appendix B.3.2 .
Figure 6: CHASEDB1. No Reg versus DualKL in PCA and latent-statistics views; full four-way comparisons are in Appendix B.3.2 .
Configuration
BAcc
Macro-F1
DualKL + simplex + shared fixed τ
72.11 ± 1.42
69.97 ± 1.71
DualKL + random + shared fixed τ
71.42 ± 1.53
69.11 ± 1.37
DualKL + empirical + shared fixed τ
71.63 ± 1.72
69.94 ± 2.08
DualKL + simplex + classwise learned τ
71.88 ± 2.62
69.65 ± 2.54
Table 2: BloodMNIST 1% training-data ablation results reported as mean ± standard deviation.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Metric
No Reg
CondKL
StdKL
DualKL
Standard
BAcc
92.90 ± 0.34
94.36 ± 0.53
94.35 ± 0.17
94.52 ± 0.22
Macro-F1
93.14 ± 0.26
94.51 ± 0.50
94.53 ± 0.18
94.72 ± 0.19
1% Training Data
BAcc
63.49 ± 1.07
71.84 ± 1.78
68.11 ± 2.65
72.22 ± 1.38
Macro-F1
61.43 ± 0.74
69.97 ± 1.83
66.25 ± 3.01
69.98 ± 1.85
4-shot
BAcc
45.31 ± 1.77
46.50 ± 3.09
47.10 ± 2.97
48.72 ± 2.59
Macro-F1
43.19 ± 1.97
44.47 ± 2.90
45.02 ± 2.44
46.57 ± 2.58
Appendix
Table 3: BloodMNIST classification results reported as mean ± standard deviation.
Setting
Metric
Strong WD
Head Dropout
Label Smoothing
Center Loss
DualKL
Standard
BAcc
93.02 ± 0.35
93.08 ± 0.63
93.05 ± 0.35
93.81 ± 0.46
94.70 ± 0.30
Macro-F1
93.11 ± 0.39
93.23 ± 0.64
93.33 ± 0.32
93.95 ± 0.43
94.88 ± 0.23
1% Training Data
BAcc
62.21 ± 1.26
63.32 ± 1.45
65.05 ± 0.97
68.85 ± 1.56
71.29 ± 1.57
Macro-F1
60.40 ± 1.53
61.79 ± 1.78
62.85 ± 1.07
66.99 ± 1.23
68.98 ± 1.36
4-shot
BAcc
43.88 ± 1.36
45.22 ± 2.74
45.74 ± 3.46
46.80 ± 1.95
48.01 ± 2.95
Macro-F1
42.08 ± 1.10
42.51 ± 2.60
43.75 ± 3.22
44.72 ± 2.14
45.91 ± 2.98
Appendix
Table 4: External baseline comparison on BloodMNIST. Results are reported as mean ± standard deviation.
Setting
Metric
No Reg
CondKL
StdKL
DualKL
4% dataset
BAcc
56.01 ± 0.54
57.79 ± 0.31
57.95 ± 0.36
58.04 ± 0.19
Macro-F1
55.87 ± 0.52
57.43 ± 0.32
57.62 ± 0.38
57.76 ± 0.19
4-shot
BAcc
28.86 ± 0.92
30.04 ± 0.90
30.12 ± 1.02
30.18 ± 1.27
Macro-F1
28.79 ± 0.80
29.45 ± 0.85
29.52 ± 0.96
29.63 ± 1.24
Appendix
Table 5: CIFAR-100 low-data classification results. We report mean ± standard deviation over seeds.
Setting
Metric
Strong WD
Head Dropout
Label Smoothing
Center Loss
DualKL
4% dataset
BAcc
55.54 ± 0.45
55.50 ± 0.27
57.14 ± 0.12
57.98 ± 0.37
58.14 ± 0.14
Macro-F1
55.44 ± 0.51
55.35 ± 0.25
56.73 ± 0.19
57.80 ± 0.35
57.86 ± 0.17
4-shot
BAcc
28.94 ± 1.03
29.11 ± 1.07
29.62 ± 1.47
30.50 ± 1.19
30.51 ± 1.37
Macro-F1
28.81 ± 0.89
29.02 ± 1.01
28.91 ± 1.46
29.96 ± 1.13
29.94 ± 1.35
Appendix
Table 6: External baseline comparison on CIFAR-100 low-data classification. Results are reported as mean ± standard deviation over evaluated seeds.
Configuration
BAcc
Macro-F1
DualKL + simplex + shared fixed τ
48.66 ± 1.99
46.59 ± 2.25
DualKL + random + shared fixed τ
46.91 ± 2.40
44.51 ± 2.82
DualKL + empirical + shared fixed τ
46.97 ± 2.83
45.19 ± 0.70
DualKL + simplex + classwise learned τ
46.43 ± 1.39
44.55 ± 0.38
Appendix
Table 7: 4-shot BloodMNIST ablation results reported as mean ± standard deviation.
Configuration
BAcc
Macro-F1
DualKL + simplex + shared fixed τ
94.74 ± 0.28
94.86 ± 0.12
DualKL + random + shared fixed τ
94.56 ± 0.24
94.63 ± 0.11
DualKL + empirical + shared fixed τ
94.52 ± 0.16
94.62 ± 0.16
DualKL + simplex + classwise learned τ
94.68 ± 0.35
94.79 ± 0.32
Appendix
Table 8: Full-training BloodMNIST ablation results reported as mean ± standard deviation.
Dataset
Metric
No Reg
StdKL
CondKL
DualKL
DRIVE
IoU
0.5558 ± 0.0061
0.5823 ± 0.0037
0.5820 ± 0.0044
0.5826 ± 0.0053
Dice
0.7143 ± 0.0051
0.7359 ± 0.0030
0.7356 ± 0.0035
0.7362 ± 0.0042
STARE
IoU
0.5049 ± 0.0311
0.5302 ± 0.0121
0.5312 ± 0.0156
0.5318 ± 0.0145
Dice
0.6701 ± 0.0278
0.6926 ± 0.0104
0.6934 ± 0.0134
0.6939 ± 0.0125
CHASEDB1
IoU
0.5525 ± 0.0066
0.5724 ± 0.0043
0.5720 ± 0.0034
0.5739 ± 0.0044
Dice
0.7116 ± 0.0054
0.7280 ± 0.0035
0.7276 ± 0.0028
0.7291 ± 0.0036
Appendix
Table 9: Repeated-seed retinal vessel segmentation results reported as mean ± standard deviation.
Setting
Metric
No Reg
StdKL
CondKL
DualKL
Full
IoU
0.8131 ± 0.0008
0.8149 ± 0.0040
0.8118 ± 0.0051
0.8158 ± 0.0017
Dice
0.8957 ± 0.0005
0.8970 ± 0.0032
0.8949 ± 0.0031
0.8976 ± 0.0011
8-shot
IoU
0.6009 ± 0.0145
0.6039 ± 0.0141
0.5439 ± 0.0471
0.6106 ± 0.0111
Dice
0.7489 ± 0.0116
0.7506 ± 0.0117
0.7015 ± 0.0394
0.7562 ± 0.0087
4-shot
IoU
0.5368 ± 0.0248
0.5460 ± 0.0221
0.5440 ± 0.0202
0.5478 ± 0.0220
Dice
0.6958 ± 0.0218
0.7036 ± 0.0188
0.7019 ± 0.0175
0.7049 ± 0.0190
Appendix
Table 10: GLAS segmentation results in full-data and few-shot regimes, reported as mean ± standard deviation.
Figure 7: BloodMNIST (standard), PCA latent geometry. Representative panel numbers: No Reg (BAcc 0.926 , Within 709.3 , Ratio 0.4369 ), DualKL (BAcc 0.952 , Within 3.4 , Ratio 0.0488 ).
Figure 13: STARE, PCA latent geometry. Representative panel numbers: No Reg (IoU 0.506 , Ratio 0.5400 ), DualKL (IoU 0.538 , Ratio 0.3840 ).
Figure 14: STARE, latent-statistics geometry. Panel medians show the No Reg to KL-regularized transition: inter-sample 0.329→0.231 – 0.237 , label-sep 0.904→2.748 – 2.857 , feat-label 0.562→0.723 – 0.731 , sKEI 1.3→4.4 – 5.2 .
Partition
Definition
K
Interpretation
Label
Cu=Yu
2
Direct foreground/background pixel label.
Count
Cu=∑v∈N(u)Yv
10
Coarse local occupancy and boundary-proximity proxy.
Pattern
Cu=bin(YN(u))
512
Exact local shape, boundary, and thin-structure pattern.
Appendix
Table 11: Segmentation condition constructions. For each pixel u , the condition Cu is induced from the ground-truth mask in a 3×3 neighborhood N(u) .
Dataset
Label K=2
Count K=10
Top-1
Hnorm
Top-1
Hnorm
CHASEDB1
92.9
0.37
81.3
0.37
DRIVE
91.4
0.42
75.8
0.44
GLAS
50.3
1.00
46.5
0.45
STARE
92.7
0.38
80.6
0.38
Appendix
Table 12: Train-split concentration statistics for coarse segmentation conditions. Top-1 mass is reported in percentage; Hnorm denotes normalized entropy.
Dataset
Obs. K
Nontriv. K
BG mass
FG mass
Nontriv. mass
CHASEDB1
489
487
81.3
0.3
18.4
DRIVE
508
506
75.8
0.3
23.9
GLAS
269
267
46.5
46.4
7.2
STARE
480
478
80.6
0.4
19.0
Appendix
Table 13: Train-split statistics for the exact 3×3 pattern partition ( K=512 ). Nontrivial patterns exclude pure-background and pure-foreground neighborhoods. Mass values are percentages.
Dataset
Label K=2
Count K=10
Pattern K=512
CHASEDB1
56.75 / 72.38
57.58 / 73.06
57.16 / 72.73
DRIVE
58.27 / 73.63
58.29 / 73.64
58.13 / 73.51
GLAS
81.83 / 89.89
81.93 / 89.93
82.25 / 90.18
STARE
52.77 / 69.06
52.85 / 69.13
53.47 / 69.67
Average
62.40 / 76.24
62.66 / 76.44
62.75 / 76.52
Appendix
Table 14: Partition-choice ablation for segmentation. We compare the direct pixel label ( K=2 ), 3×3 foreground-count condition ( K=10 ), and exact binary local-mask pattern ( K=512 ). Results are validation IoU/Dice (%).
Figure 15: Sensitivity to (βstd,βcond) on full-data BloodMNIST. Each cell reports the change relative to the grid mean.
Related method
How our method differs
Deep VIB [ 3 ]
Deep VIB provides a variational KL surrogate for a single global information bottleneck, mainly controlling I(Z;X) . Our method keeps this global control through Jstd , but adds a label-induced conditional control axis Jcond motivated by I(Z;X)=I(Z;C)+I(Z;X∣C) . Thus, our objective separates global capacity control from within-condition information control.
Multi-view conditional VIB-style objectives [ 7 ]
These objectives use conditional information terms to calibrate shared representations across multiple observed views and to handle view imbalance or missing views. Our conditioning variable is not a view but a supervised label-induced partition C=π(Y) , and our goal is not multi-view calibration but supervised regularization of within-condition residual information.
This line of work uses conditional information bottleneck-style objectives to compress self-supervised representations, typically between augmented views. Our method uses a supervised condition variable C=π(Y) and combines conditional information control with an explicit global KL term. Hence, the method is a supervised dual-bottleneck regularizer rather than a self-supervised compression objective.
Conditional VAE / deep conditional generative models [ 11 ]
Conditional VAEs use condition-dependent latent variables or priors inside a generative likelihood or ELBO for structured output modelling. Our conditional prior is not used for generation; it is a variational information-control prior inside a supervised regularizer, where the conditional KL decomposes into I(Z;X∣C) plus condition-wise prior mismatch.
Semi-supervised deep generative models [ 12 ]
Semi-supervised generative models use class variables and latent variables to model data density and exploit unlabeled data. Our method does not model p(x) or require unlabeled data; it regularizes a supervised task representation by jointly controlling global capacity and within-condition information.
Appendix
Table 15: Comparison with related bottleneck and conditional-prior formulations.
Research Center Trustworthy Data Science and Security, Dortmund · Faculty of Computer Science, Ruhr University Bochum · Signal Processing and Speech Communication Laboratory, Graz University of Technology +2