Rethinking the Information Bottleneck: Structured Decomposition under Label-Induced Partitions
Organizations: School of Computer Science, The University of Sydney Sydney, NSW 2006, Australia
Abstract
Standard information bottleneck (IB) regularization constrains representations via a single scalar I(Z;X), implicitlytreating all information as homogeneous. However, a single global compression control couples label-relevant structurewith residual within-condition variation, rather than regulating their allocation independently, allowing nuisanceinformation to persist in learned representations. For example, in medical imaging applications, residual variation oftenstems from acquisition conditions, background factors, or subject-specific appearance. This issue becomes particularlypronounced in data-limited settings, where models tend to overfit such variation, hindering generalization. While existingregularization methods can stabilize training, control capacity, or shape representation geometry, they do not explicitlyseparate nuisance-like variation from task-supporting structure. To address this limitation, we revisit IB from a structuredperspective based on a label-induced partition, where condition-level structure and within-condition information playdistinct roles. This leads to a dual-bottleneck formulation: a standard KL term controls global information capacity, while aconditional KL term targets within-condition information. We show that the conditional KL admits an exact decompositioninto a within-condition information term and a prior-mismatch term, explaining its alignment with the design objective.With a simplex-structured conditional prior, the method provides controllable latent geometry and integrates seamlesslyinto existing pipelines. Experiments on classification and segmentation show the clearest gains in low-data classificationand consistent improvements across dense prediction benchmarks.
Figures & tables
| Method | Global | Cond. | Geometry / prior | Main distinction |
|---|---|---|---|---|
| VIB / StdKL | – | global prior | only global bottleneck | |
| CondKL (cond-only) | – | conditional prior | no label-induced dual control | |
| Center / prototype | – | implicit | empirical centers | feature-space geometry |
| SupCon | – | implicit | contrastive pairs | pairwise geometry |
| Generic regularizers | generic | – | none | orthogonal heuristics |
| DualKL (ours) | simplex conditional prior | joint two-axis control |
| Configuration | BAcc | Macro-F1 |
|---|---|---|
| DualKL + simplex + shared fixed | 72.11 1.42 | 69.97 1.71 |
| DualKL + random + shared fixed | 71.42 1.53 | 69.11 1.37 |
| DualKL + empirical + shared fixed | 71.63 1.72 | 69.94 2.08 |
| DualKL + simplex + classwise learned | 71.88 2.62 | 69.65 2.54 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | Metric | No Reg | CondKL | StdKL | DualKL |
|---|---|---|---|---|---|
| Standard | BAcc | 92.90 0.34 | 94.36 0.53 | 94.35 0.17 | 94.52 0.22 |
| Macro-F1 | 93.14 0.26 | 94.51 0.50 | 94.53 0.18 | 94.72 0.19 | |
| 1% Training Data | BAcc | 63.49 1.07 | 71.84 1.78 | 68.11 2.65 | 72.22 1.38 |
| Macro-F1 | 61.43 0.74 | 69.97 1.83 | 66.25 3.01 | 69.98 1.85 | |
| 4-shot | BAcc | 45.31 1.77 | 46.50 3.09 | 47.10 2.97 | 48.72 2.59 |
| Macro-F1 | 43.19 1.97 | 44.47 2.90 | 45.02 2.44 | 46.57 2.58 |
| Setting | Metric | Strong WD | Head Dropout | Label Smoothing | Center Loss | DualKL |
|---|---|---|---|---|---|---|
| Standard | BAcc | 93.02 0.35 | 93.08 0.63 | 93.05 0.35 | 93.81 0.46 | 94.70 0.30 |
| Macro-F1 | 93.11 0.39 | 93.23 0.64 | 93.33 0.32 | 93.95 0.43 | 94.88 0.23 | |
| 1% Training Data | BAcc | 62.21 1.26 | 63.32 1.45 | 65.05 0.97 | 68.85 1.56 | 71.29 1.57 |
| Macro-F1 | 60.40 1.53 | 61.79 1.78 | 62.85 1.07 | 66.99 1.23 | 68.98 1.36 | |
| 4-shot | BAcc | 43.88 1.36 | 45.22 2.74 | 45.74 3.46 | 46.80 1.95 | 48.01 2.95 |
| Macro-F1 | 42.08 1.10 | 42.51 2.60 | 43.75 3.22 | 44.72 2.14 | 45.91 2.98 |
| Setting | Metric | No Reg | CondKL | StdKL | DualKL |
|---|---|---|---|---|---|
| 4% dataset | BAcc | 56.01 0.54 | 57.79 0.31 | 57.95 0.36 | 58.04 0.19 |
| Macro-F1 | 55.87 0.52 | 57.43 0.32 | 57.62 0.38 | 57.76 0.19 | |
| 4-shot | BAcc | 28.86 0.92 | 30.04 0.90 | 30.12 1.02 | 30.18 1.27 |
| Macro-F1 | 28.79 0.80 | 29.45 0.85 | 29.52 0.96 | 29.63 1.24 |
| Setting | Metric | Strong WD | Head Dropout | Label Smoothing | Center Loss | DualKL |
|---|---|---|---|---|---|---|
| 4% dataset | BAcc | 55.54 0.45 | 55.50 0.27 | 57.14 0.12 | 57.98 0.37 | 58.14 0.14 |
| Macro-F1 | 55.44 0.51 | 55.35 0.25 | 56.73 0.19 | 57.80 0.35 | 57.86 0.17 | |
| 4-shot | BAcc | 28.94 1.03 | 29.11 1.07 | 29.62 1.47 | 30.50 1.19 | 30.51 1.37 |
| Macro-F1 | 28.81 0.89 | 29.02 1.01 | 28.91 1.46 | 29.96 1.13 | 29.94 1.35 |
| Configuration | BAcc | Macro-F1 |
|---|---|---|
| DualKL + simplex + shared fixed | 48.66 1.99 | 46.59 2.25 |
| DualKL + random + shared fixed | 46.91 2.40 | 44.51 2.82 |
| DualKL + empirical + shared fixed | 46.97 2.83 | 45.19 0.70 |
| DualKL + simplex + classwise learned | 46.43 1.39 | 44.55 0.38 |
| Configuration | BAcc | Macro-F1 |
|---|---|---|
| DualKL + simplex + shared fixed | 94.74 0.28 | 94.86 0.12 |
| DualKL + random + shared fixed | 94.56 0.24 | 94.63 0.11 |
| DualKL + empirical + shared fixed | 94.52 0.16 | 94.62 0.16 |
| DualKL + simplex + classwise learned | 94.68 0.35 | 94.79 0.32 |
| Dataset | Metric | No Reg | StdKL | CondKL | DualKL |
|---|---|---|---|---|---|
| DRIVE | IoU | 0.5558 0.0061 | 0.5823 0.0037 | 0.5820 0.0044 | 0.5826 0.0053 |
| Dice | 0.7143 0.0051 | 0.7359 0.0030 | 0.7356 0.0035 | 0.7362 0.0042 | |
| STARE | IoU | 0.5049 0.0311 | 0.5302 0.0121 | 0.5312 0.0156 | 0.5318 0.0145 |
| Dice | 0.6701 0.0278 | 0.6926 0.0104 | 0.6934 0.0134 | 0.6939 0.0125 | |
| CHASEDB1 | IoU | 0.5525 0.0066 | 0.5724 0.0043 | 0.5720 0.0034 | 0.5739 0.0044 |
| Dice | 0.7116 0.0054 | 0.7280 0.0035 | 0.7276 0.0028 | 0.7291 0.0036 |
| Setting | Metric | No Reg | StdKL | CondKL | DualKL |
|---|---|---|---|---|---|
| Full | IoU | 0.8131 0.0008 | 0.8149 0.0040 | 0.8118 0.0051 | 0.8158 0.0017 |
| Dice | 0.8957 0.0005 | 0.8970 0.0032 | 0.8949 0.0031 | 0.8976 0.0011 | |
| 8-shot | IoU | 0.6009 0.0145 | 0.6039 0.0141 | 0.5439 0.0471 | 0.6106 0.0111 |
| Dice | 0.7489 0.0116 | 0.7506 0.0117 | 0.7015 0.0394 | 0.7562 0.0087 | |
| 4-shot | IoU | 0.5368 0.0248 | 0.5460 0.0221 | 0.5440 0.0202 | 0.5478 0.0220 |
| Dice | 0.6958 0.0218 | 0.7036 0.0188 | 0.7019 0.0175 | 0.7049 0.0190 |
| Partition | Definition | Interpretation | |
|---|---|---|---|
| Label | 2 | Direct foreground/background pixel label. | |
| Count | 10 | Coarse local occupancy and boundary-proximity proxy. | |
| Pattern | 512 | Exact local shape, boundary, and thin-structure pattern. |
| Dataset | Label | Count | ||
|---|---|---|---|---|
| Top-1 | Top-1 | |||
| CHASEDB1 | 92.9 | 0.37 | 81.3 | 0.37 |
| DRIVE | 91.4 | 0.42 | 75.8 | 0.44 |
| GLAS | 50.3 | 1.00 | 46.5 | 0.45 |
| STARE | 92.7 | 0.38 | 80.6 | 0.38 |
| Dataset | Obs. | Nontriv. | BG mass | FG mass | Nontriv. mass |
|---|---|---|---|---|---|
| CHASEDB1 | 489 | 487 | 81.3 | 0.3 | 18.4 |
| DRIVE | 508 | 506 | 75.8 | 0.3 | 23.9 |
| GLAS | 269 | 267 | 46.5 | 46.4 | 7.2 |
| STARE | 480 | 478 | 80.6 | 0.4 | 19.0 |
| Dataset | Label | Count | Pattern |
|---|---|---|---|
| CHASEDB1 | 56.75 / 72.38 | 57.58 / 73.06 | 57.16 / 72.73 |
| DRIVE | 58.27 / 73.63 | 58.29 / 73.64 | 58.13 / 73.51 |
| GLAS | 81.83 / 89.89 | 81.93 / 89.93 | 82.25 / 90.18 |
| STARE | 52.77 / 69.06 | 52.85 / 69.13 | 53.47 / 69.67 |
| Average | 62.40 / 76.24 | 62.66 / 76.44 | 62.75 / 76.52 |
| Related method | How our method differs |
|---|---|
| Deep VIB [ 3 ] | Deep VIB provides a variational KL surrogate for a single global information bottleneck, mainly controlling . Our method keeps this global control through , but adds a label-induced conditional control axis motivated by . Thus, our objective separates global capacity control from within-condition information control. |
| Multi-view conditional VIB-style objectives [ 7 ] | These objectives use conditional information terms to calibrate shared representations across multiple observed views and to handle view imbalance or missing views. Our conditioning variable is not a view but a supervised label-induced partition , and our goal is not multi-view calibration but supervised regularization of within-condition residual information. |
| Self-supervised compressive visual representations [ 8 ] | This line of work uses conditional information bottleneck-style objectives to compress self-supervised representations, typically between augmented views. Our method uses a supervised condition variable and combines conditional information control with an explicit global KL term. Hence, the method is a supervised dual-bottleneck regularizer rather than a self-supervised compression objective. |
| Conditional VAE / deep conditional generative models [ 11 ] | Conditional VAEs use condition-dependent latent variables or priors inside a generative likelihood or ELBO for structured output modelling. Our conditional prior is not used for generation; it is a variational information-control prior inside a supervised regularizer, where the conditional KL decomposes into plus condition-wise prior mismatch. |
| Semi-supervised deep generative models [ 12 ] | Semi-supervised generative models use class variables and latent variables to model data density and exploit unlabeled data. Our method does not model or require unlabeled data; it regularizes a supervised task representation by jointly controlling global capacity and within-condition information. |