Anatomy-Structured Hierarchical MIL for Weakly-Supervised Thoracic Disease Detection in Chest X-rays
Authors: Jeongin Kim, Sohyun Ahn, Seo Young Kang, Jaeyi Sung, Soomin Kim, Sungho Cho, Rena Lee, Kwanchang Kim, +1 more
Organizations: Division of Artificial Intelligence & Software, Ewha Womans University · Ewha Medical Artificial Intelligence Research Institute, Ewha Womans University · Department of Nuclear Medicine, Ewha Womans University · REMEDI Inc. R&D Center · Ewha Womans University Seoul Hospital
Weakly-supervised thoracic disease detection in chest X-rays (CXR) is challenging due to subtle appearances and complex anatomical overlap, motivating anatomy-aware modeling for improved localization. However, prior anatomy-aware methods typically rely on coarse region proxies or static spatial priors, which may restrict dynamic instance discovery and limit precise localization of small abnormalities. We propose Anatomy-Structured Hierarchical Multiple Instance Learning (ASH-MIL), a framework that introduces parallel anatomy-structured observation branches (cardiac, pulmonary, and agnostic) combined with hierarchical MIL aggregation. Anatomical priors are injected as soft spatial biases into decoder cross-attention, enabling anatomically grounded evidence maps without disease bounding-box supervision. Instance localization is derived directly from MIL-weighted cross-attention maps without bounding box supervision. Experiments on CXR8 and cross-domain MIMIC-CXR held-out sets demonstrate consistent improvements over prior weakly-supervised and anatomy-aware approaches, particularly under stricter localization criteria. Our code is available at https://github.com/jn-kim/ash-mil.
Figures & tables
Figure 1: Overall architecture of ASH-MIL. A shared RAD-DINO backbone extracts patch tokens that are processed by three parallel anatomy-structured decoder branches: cardiac ( car ), pulmonary ( pul ), and agnostic ( agn ). For brevity, only the pulmonary branch is illustrated in detail. The right panel presents the hierarchical MIL framework that aggregates query-level and branch-level evidence to produce the final class-wise predictions.
IoU@0.1
IoU@0.3
IoU@0.5
IoU@0.1–0.5
Method
AP
CorLoc
AP
CorLoc
AP
CorLoc
mAP
CorLoc
CXR8
ASH-MIL (Ours)
28.82
0.81
14.23
0.59
6.03
0.26
15.77
0.55
Ablations
3 branch (w/o prior)
25.02
0.68
11.48
0.44
2.05
0.18
13.07
0.43
1 branch (w/ prior)
22.55
0.60
8.86
0.32
3.31
0.13
11.35
0.34
Table 1: Detection performance on CXR8 and MIMIC-CXR datasets.
Figure 2: Qualitative results on CXR8. Top: the input image and branch-specific anatomical priors (cardiac, pulmonary, agnostic). Bottom: the final Nodule localization (GT: red; Pred: green) and the corresponding branch evidence maps. The learned branch attention weights αp quantify each branch’s contribution to the Nodule prediction.
Vision-language models (VLMs) pretrained on large-scale image-text pairs demonstrate strong image-level understanding, but are primarily optimized for global alignment and do not explicitly encode fine-grained anatomical structure, limiting their suitability for spatially precise tasks such as segmentation. We introduce CheXanatomy, a framework that integrates explicit anatomical knowledge into a pretrained VLM through autoregressive token-space supervision. Instead of adding task-specific decoder heads, the model is trained to generate anatomical segmentation masks via next-token prediction. To enable scalable supervision, we synthesize realistic chest radiographs from CT volumes and forward-project CT segmentation labels to obtain anatomically consistent 2D masks. We evaluate the approach on synthetic and real chest radiographs against a U-Net baseline, including ablations on model scale, input resolution, and vision encoder fine-tuning. Autoregressive anatomical supervision achieves performance comparable to specialized convolutional models in-distribution and demonstrates improved geometric robustness under domain shift to real CXR data. In addition, anatomy-pretrained models exhibit improved sample efficiency when adapting to novel localization tasks under limited supervision. Larger models and higher input image resolution improve performance, while vision encoder fine-tuning has limited effect. These results show that embedding anatomical structure directly into the generative objective promotes spatially grounded representations and supports anatomy-aware medical vision-language modeling.
Sergios Gatidis, Curtis Langlotz, Christian Bluethgen
Stanford Center for Artificial Intelligence in Medicine and Imaging, Stanford University, Palo Alto, CA, USA · Department of Radiology, Stanford University, Stanford, CA, USA
Accurate chest X-ray interpretation is inherently hierarchical. Clinical decisions depend not only on what abnormality is present but where it is situated, requiring reasoning from broad anatomical systems down to specific pathological findings. Yet existing automated systems largely treat this as a flat classification problem, failing to capture inter-level dependencies or enforce coherence between coarse and fine predictions. We propose CHASE (Classification with Hierarchical Analysis and Structured Enforcement), a unified single-stage framework that mirrors radiologists' coarse-to-fine reasoning through a clinically driven three-level taxonomy of 9 anatomical regions, 17 sub-regions, and 28 pathological findings. CHASE jointly optimizes multi-level supervision, cross-level probability alignment, and a hierarchy-violation penalty within a shared Vision Transformer backbone. This ensures that fine-grained findings are anatomically supported by their coarser-level context rather than predicted in isolation. Experiments demonstrate that CHASE outperforms flat and hierarchical baselines across all levels while achieving superior probabilistic hierarchy consistency, with level-wise attention maps confirming anatomically grounded predictions. Code is available at: https://github.com/yejix-ai/CHASE.
Deep learning models for CT scan analysis are often limited by the scarcity of precise pixel-level annotations, which require significant radiologist effort to produce. Training on scan-level labels alone reduces annotation requirements but introduces challenges: low supervision ratios and large input volumes make models prone to overfitting and shortcut learning. In this work, we investigate two complementary methods to address these challenges: multi-instance learning (MIL) and anatomical filtering. MIL divides CT volumes into 2D slice instances, enabling efficient 2D architectures with ImageNet pretraining rather than computationally demanding 3D models. Anatomical filtering uses Compass, our self-supervised body part regression model, to crop scans to pathology-relevant subregions without requiring segmentation masks. We evaluate two MIL frameworks - Attention-based MIL (ABMIL) and FocusMIL - on kidney tumor classification across one internal dataset (TUH) and two external datasets (KiTS23 and TCGA-KiRC). Our best models achieve F1 = 0.83 on the internal test set using only scan-level labels. We further show that anatomical filtering with the Compass model is critical for the out-of-distribution generalization of embedding-based ABMIL, while instance-based FocusMIL demonstrates greater inherent robustness to distribution shift. While evaluated on kidney tumors, we consider this a proof-of-concept for a broader weakly supervised CT classification pipeline applicable to other organs and pathologies.
Joonas Ariva, Dmytro Fishman
University of Tartu, Tartu, Estonia · STACC, Tartu, Estonia · Better Medicine, Tartu, Estonia