Quasi-Binarized Autoencoders: An Architecture-Independent Information Bottleneck for Medical Image Anomaly Detection
Organizations: Department of Radiology, The University of Tokyo Hospital, 7-3-1 Hongo, Bunkyo-ku, Tokyo, 113-8655, Japan · Department of Computational Diagnostic Radiology and Preventive Medicine, The University of Tokyo Hospital, 7-3-1 Hongo, Bunkyo-ku, Tokyo, 113-8655, Japan · Department of Radiology, Kanazawa University Hospital, 13-1 Takara-machi, Kanazawa, Ishikawa, 920-8641, Japan
Abstract
Unsupervised anomaly detection, which learns only from normal images, is a central task in medical image analysis and remains an open problem. Reconstruction-based methods pass an image through an encoder-decoder network trained on normal data and detect anomalies from the residual between the image and its reconstruction. This works only if the information passed from the encoder to the decoder is limited; otherwise the network learns an identity mapping and reconstructs anomalies too. This limit is usually imposed through architectural choices, tuned per dataset, that cannot be stated in bits. We introduce the quasi-binarizing (QB) layer, which squashes each latent element into [0, 1] and adds Laplace noise of scale 1/epsilon. Each element is then epsilon-locally differentially private, and the mutual information between an image and its reconstruction is bounded by a quantity that depends only on epsilon and the number of QB elements, whatever the encoder and decoder. Placing a QB layer on every encoder-decoder path, including all skip connections, we build QBAE, a seven-level attention U-Net with 32,768 QB elements. On the seven datasets of the MedIAnomaly benchmark, QBAE with one architecture and one configuration reaches a mean image-level AUROC of 0.828, the highest among methods that do not adapt to each dataset, and the best reported results on BraTS2021 (AUROC 0.911, pixel-level AP 0.838). The noise is kept at test time, so that every reconstruction satisfies the bound. Without input corruption, the bottleneck alone prevents identity collapse (mean AUROC 0.805 vs. 0.590). Code is available at https://github.com/hanaokalog/MedIAnomalyQB.
Figures & tables
| Setting | Per element | Total budget | Relative to a raw 8-bit image |
|---|---|---|---|
| 0.11 bits | bits | 0.03 | |
| 0.19 bits | bits | 0.05 | |
| 0.65 bits | bits | 0.16 | |
| 1.98 bits | bits | 0.50 | |
| 3.52 bits | bits | 0.88 | |
| 5.25 bits | bits | 1.31 |
| Level | Resolution | Channels per pixel | QB elements | Share of |
|---|---|---|---|---|
| 1 | 1 | 16,384 | 50.0% | |
| 2 | 2 | 8,192 | 25.0% | |
| 3 | 4 | 4,096 | 12.5% | |
| 4 | 8 | 2,048 | 6.3% | |
| 5 | 16 | 1,024 | 3.1% | |
| 6 | 32 | 512 | 1.6% |
| Dataset | Source | Modality | Train (normal) | Test normal | Test abnormal | Pixel labels |
|---|---|---|---|---|---|---|
| RSNA | Shih et al. (2019) a | Chest X-ray | 3,851 | 1,000 | 1,000 | – |
| VinDr-CXR | Nguyen et al. (2022) | Chest X-ray | 4,000 | 1,000 | 1,000 | – |
| Brain Tumor | Hamada (2025); Saleh et al. (2020); Cheng et al. (2015) b | Brain MRI | 1,000 | 600 | 600 | – |
| LAG | Li et al. (2019) | Fundus photograph | 1,500 | 811 | 811 | – |
| ISIC2018 | Codella et al. (2019) | Dermoscopy | 6,705 | 909 | 603 | – |
| Camelyon16 | Ehteshami Bejnordi et al. (2017); Bao et al. (2024) c | Histopathology | 5,088 | 1,120 | 1,113 | – |
| Dataset | QBAE AUROC | QBAE AP | AE-PL | DAE | MSDE AUROC (AP) | Best in MedIAnomaly |
|---|---|---|---|---|---|---|
| RSNA | 0.881 0.002 | 0.858 0.005 | 0.875 | 0.861 | 0.918 (0.906) | 0.911 (FAE-SSIM † ) |
| VinDr-CXR | 0.723 0.009 | 0.730 0.010 | 0.753 | 0.686 | 0.819 (0.797) | 0.769 (AE-U) |
| Brain Tumor | 0.959 0.003 | 0.920 0.006 | 0.957 | 0.832 | 0.981 (0.981) | 0.973 (MSC † ) |
| LAG | 0.844 0.010 | 0.773 0.011 | 0.856 | 0.716 | 0.810 (0.831) | 0.856 (AE-PL † ) |
| ISIC2018 | 0.751 0.009 | 0.656 0.009 | 0.684 | 0.700 | 0.705 (0.638) | 0.807 (AnatPaste) |
| Camelyon16 | 0.726 0.005 | 0.647 0.006 | 0.761 | 0.654 | 0.812 (0.820) | 0.807 (AutoDDPM) |
| Setting | Image AUROC | Pixel AP | Best Dice |
|---|---|---|---|
| Common setting ( = 10, KL on) | 0.911 0.017 | 0.838 0.016 | 0.773 0.017 |
| = 30, KL on | 0.939 0.003 | 0.849 0.008 | 0.786 0.006 |
| = 30, KL off (best on BraTS) | 0.942 0.003 | 0.855 0.002 | 0.792 0.003 |
| DAE (best in MedIAnomaly) | 0.859 | 0.755 | 0.711 |
| Dataset | QB off | = 10 (common) | Best | Best of all settings |
|---|---|---|---|---|
| RSNA | 0.832 | 0.881 | 0.890 ( = 3) | 0.893 ( = 3, corr. off, KL on) |
| VinDr-CXR | 0.734 | 0.723 | 0.748 ( = 30) | 0.753 ( = 100, corr. on, KL off) |
| Brain Tumor | 0.961 | 0.959 | 0.965 ( = 30) | 0.968 ( = off, corr. on, KL off) |
| LAG | 0.802 | 0.844 | 0.844 ( = 10) | 0.860 ( = 3, corr. on, KL off) |
| ISIC2018 | 0.738 | 0.751 | 0.751 ( = 10) | 0.752 ( = 30, corr. on, KL off) |
| Camelyon16 | 0.593 | 0.726 | 0.794 ( = 1) | 0.803 ( = 1, corr. off, KL on) |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| [nats] | [nats] | Bound per element [bits] | Total [bits] | |
|---|---|---|---|---|
| 0.3 | 0.09 | 0.078 | 0.112 | 3,685 |
| 1 | 1 | 0.131 | 0.189 | 6,205 |
| 3 | 3 | 0.449 | 0.648 | 21,238 |
| 10 | 10 | 1.374 | 1.982 | 64,941 |
| 30 | 30 | 2.438 | 3.518 | 115,267 |
| 100 | 100 | 3.638 | 5.249 | 171,994 |
| Level | Resolution | Width | QB elements (channels per pixel) | Decoder attention | Parameters (enc. / dec.) |
|---|---|---|---|---|---|
| 1 | 16 | 16,384 (1) | – | 0.003 M / 0.03 M | |
| 2 | 32 | 8,192 (2) | channel attention + FFN | 0.015 M / 0.12 M | |
| 3 | 64 | 4,096 (4) | spatial attention + FFN | 0.06 M / 0.46 M | |
| 4 | 128 | 2,048 (8) | spatial attention + FFN | 0.23 M / 1.84 M | |
| 5 | 256 | 1,024 (16) | spatial attention + FFN | 0.92 M / 7.35 M | |
| 6 | 512 | 512 (32) | spatial attention + FFN | 3.67 M / 29.39 M |