cs.CVSep 27, 2026

Anatomy-Structured Hierarchical MIL for Weakly-Supervised Thoracic Disease Detection in Chest X-rays

Authors: Jeongin Kim, Sohyun Ahn, Seo Young Kang, Jaeyi Sung, Soomin Kim, Sungho Cho, Rena Lee, Kwanchang Kim, +1 more

Organizations: Division of Artificial Intelligence & Software, Ewha Womans University · Ewha Medical Artificial Intelligence Research Institute, Ewha Womans University · Department of Nuclear Medicine, Ewha Womans University · REMEDI Inc. R&D Center · Ewha Womans University Seoul Hospital

Abstract

Weakly-supervised thoracic disease detection in chest X-rays (CXR) is challenging due to subtle appearances and complex anatomical overlap, motivating anatomy-aware modeling for improved localization. However, prior anatomy-aware methods typically rely on coarse region proxies or static spatial priors, which may restrict dynamic instance discovery and limit precise localization of small abnormalities. We propose Anatomy-Structured Hierarchical Multiple Instance Learning (ASH-MIL), a framework that introduces parallel anatomy-structured observation branches (cardiac, pulmonary, and agnostic) combined with hierarchical MIL aggregation. Anatomical priors are injected as soft spatial biases into decoder cross-attention, enabling anatomically grounded evidence maps without disease bounding-box supervision. Instance localization is derived directly from MIL-weighted cross-attention maps without bounding box supervision. Experiments on CXR8 and cross-domain MIMIC-CXR held-out sets demonstrate consistent improvements over prior weakly-supervised and anatomy-aware approaches, particularly under stricter localization criteria. Our code is available at https://github.com/jn-kim/ash-mil.

Figures & tables

Explore similar work

Jun 7, 2026cs.CV

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

Vision-language models (VLMs) pretrained on large-scale image-text pairs demonstrate strong image-level understanding, but are primarily optimized for global alignment and do not explicitly encode fine-grained anatomical structure, limiting their suitability for spatially precise tasks such as segmentation. We introduce CheXanatomy, a framework that integrates explicit anatomical knowledge into a pretrained VLM through autoregressive token-space supervision. Instead of adding task-specific decoder heads, the model is trained to generate anatomical segmentation masks via next-token prediction. To enable scalable supervision, we synthesize realistic chest radiographs from CT volumes and forward-project CT segmentation labels to obtain anatomically consistent 2D masks. We evaluate the approach on synthetic and real chest radiographs against a U-Net baseline, including ablations on model scale, input resolution, and vision encoder fine-tuning. Autoregressive anatomical supervision achieves performance comparable to specialized convolutional models in-distribution and demonstrates improved geometric robustness under domain shift to real CXR data. In addition, anatomy-pretrained models exhibit improved sample efficiency when adapting to novel localization tasks under limited supervision. Larger models and higher input image resolution improve performance, while vision encoder fine-tuning has limited effect. These results show that embedding anatomical structure directly into the generative objective promotes spatially grounded representations and supports anatomy-aware medical vision-language modeling.
Aug 4, 2026cs.CV

Clinically-Grounded Hierarchical Classification for Consistent Chest X-ray Interpretation

Accurate chest X-ray interpretation is inherently hierarchical. Clinical decisions depend not only on what abnormality is present but where it is situated, requiring reasoning from broad anatomical systems down to specific pathological findings. Yet existing automated systems largely treat this as a flat classification problem, failing to capture inter-level dependencies or enforce coherence between coarse and fine predictions. We propose CHASE (Classification with Hierarchical Analysis and Structured Enforcement), a unified single-stage framework that mirrors radiologists' coarse-to-fine reasoning through a clinically driven three-level taxonomy of 9 anatomical regions, 17 sub-regions, and 28 pathological findings. CHASE jointly optimizes multi-level supervision, cross-level probability alignment, and a hierarchy-violation penalty within a shared Vision Transformer backbone. This ensures that fine-grained findings are anatomically supported by their coarser-level context rather than predicted in isolation. Experiments demonstrate that CHASE outperforms flat and hierarchical baselines across all levels while achieving superior probabilistic hierarchy consistency, with level-wise attention maps confirming anatomically grounded predictions. Code is available at: https://github.com/yejix-ai/CHASE.
Sep 7, 2026cs.CV

Weakly-supervised Kidney Tumor Classification from CT Scans with Multi-Instance Learning and Anatomical Filtering

Deep learning models for CT scan analysis are often limited by the scarcity of precise pixel-level annotations, which require significant radiologist effort to produce. Training on scan-level labels alone reduces annotation requirements but introduces challenges: low supervision ratios and large input volumes make models prone to overfitting and shortcut learning. In this work, we investigate two complementary methods to address these challenges: multi-instance learning (MIL) and anatomical filtering. MIL divides CT volumes into 2D slice instances, enabling efficient 2D architectures with ImageNet pretraining rather than computationally demanding 3D models. Anatomical filtering uses Compass, our self-supervised body part regression model, to crop scans to pathology-relevant subregions without requiring segmentation masks. We evaluate two MIL frameworks - Attention-based MIL (ABMIL) and FocusMIL - on kidney tumor classification across one internal dataset (TUH) and two external datasets (KiTS23 and TCGA-KiRC). Our best models achieve F1 = 0.83 on the internal test set using only scan-level labels. We further show that anatomical filtering with the Compass model is critical for the out-of-distribution generalization of embedding-based ABMIL, while instance-based FocusMIL demonstrates greater inherent robustness to distribution shift. While evaluated on kidney tumors, we consider this a proof-of-concept for a broader weakly supervised CT classification pipeline applicable to other organs and pathologies.