cs.CVSep 28, 2026

XFlow: A Workflow Model for Instruction-Guided Lesion Segmentation in Chest X-rays

Authors: Geon Choi, Hangyul Yoon, Hyunki Park, Sang Hoon Seo, Edward Choi

Organizations: KAIST · Korea University Guro Hospital · Samsung Medical Center

Abstract

Existing text-guided segmentation models in the medical domain cover only a narrow set of anatomical structures and lesions in chest X-rays (CXRs), and most of them assume that the queried target is always present in the image. Instruction-guided lesion segmentation (ILS) was introduced to overcome these limitations by segmenting diverse lesion types from simple user instructions while also recognizing when the queried lesion is absent, and ROSALIA was proposed as the first model for this task. However, the masks produced by ROSALIA remain of limited quality, often carrying scattered noise. Moreover, ROSALIA predicts the mask in a single shot, which differs fundamentally from how radiologists perceive and delineate lesions in practice. A radiologist first surveys the entire thorax, then localizes the approximate region of abnormality, and only then refines the lesion contour. Motivated by this coarse-to-fine, multi-level perception process, we present XFlow, a workflow model for ILS that combines box-based localization with multi-turn point refinement. XFlow detects the lungs, decides whether the queried finding is present in each of them, and prompts a fine-tuned SAM with the lesion box for an initial mask. It then corrects that mask through point prompts until its boundary follows the lesion, leaving every intermediate decision visible. Our experiments show that XFlow achieves the best segmentation quality on both internal and external evaluation. Notably, it surpasses ROSALIA in segmentation quality even when the two are trained on the same lesion annotations. Code and model weights will be made publicly available.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jul 8, 2026cs.CV

Seeing What Matters: Lesion-Aware High-Resolution Patch Discovery and Fusion for Chest X-ray Report Generation

Despite rapid advances in chest X-ray (CXR) foundation models, most radiology report generation (RRG) systems still rely on heavily downsampled inputs (e.g., 256x256) due to the fixed visual token budgets of pretrained vision encoders, suppressing subtle yet clinically important cues present in native-resolution images. However, enabling high-resolution (high-res) perception remains challenging: naive tiling causes prohibitive token inflation, while global compression suppresses subtle lesions and degrades diagnostic fidelity. Inspired by radiologists' workflow, localizing suspicious regions before detailed high-res assessment. We propose Lesion-Aware High-Resolution Patch Discovery and Fusion for Chest X-ray Reporting (LePaX), the first RRG framework that enables efficient high-res CXR perception (up to 1920x1920) without increasing the vision-token count. LePaX formulates high-res perception as a constrained spatial resolution allocation problem under a fixed token budget and introduces two key components: Learnable Spatial Resolution Allocation (LSRA), which learns a spatial utility map that adaptively allocates limited high-res capacity to diagnostically relevant regions, enabling targeted extraction of high-res patches from native CXRs; and Global-Regional Fusion (GRF), which performs token-preserving region-to-global refinement by projecting high-resolution regional evidence back onto the global feature grid through spatially aligned resolution write-back, avoiding token inflation. Experiments on multiple CXR benchmarks demonstrate that LePaX consistently improves both clinical and linguistic metrics while enabling native-resolution CXR perception with over 10x fewer visual tokens than naive high-res tiling.
May 18, 2026cs.CV

Rad-VLSM: A Cross-Modal Framework with Semantics-Assisted Prompting for Medical Segmentation and Diagnosis

Medical image segmentation is more clinically valuable when it supports diagnosis rather than merely producing lesion masks. However, diagnostically relevant lesion cues are often subtle and localized, while existing models may be distracted by background tissues, acoustic artifacts, and irrelevant visual correlations. To address this problem, we propose Rad-VLSM, a two-stage cross-modal framework for semantics-assisted lesion focusing, robust segmentation, and visually grounded diagnosis. In the first stage, a BLIP-2-based vision-language alignment module identifies lesion-related candidate regions under semantic guidance and converts them into box prompts. In the second stage, these prompts are fed into a SAM-based multitask network, where a multi-candidate region aggregation strategy improves prompt stability and guides lesion segmentation. The predicted masks are then used as spatial priors for diagnosis, and a visual-radiomics fusion head integrates lesion-aware visual features with selected radiomics descriptors. By using semantic information for localization rather than direct prediction, Rad-VLSM reduces text-to-diagnosis dependence and grounds diagnosis in lesion-level evidence. Experiments on a private clinical breast ultrasound dataset and public benchmarks show that Rad-VLSM achieves strong segmentation and diagnostic performance with favorable generalization.
Jun 19, 2026cs.CV

CheXpercept: A Benchmark for Evaluating Expert-Level Lesion Perception in Chest X-rays

The evaluation of vision-language models (VLMs) for chest X-ray (CXR) analysis has largely been limited to disease-presence classification without visual grounding. Such evaluations fail to verify the expert-level lesion perception necessary to ensure the clinical reliability of VLMs. To address these limitations, we introduce CheXpercept, a sequential, multi-level perception benchmark that mirrors a radiologist's cognitive workflow across coarse-level detection, fine-level contour evaluation and revision, and semantic-level attribute extraction. To ensure high clinical fidelity at scale, we construct the dataset using a semi-automated generation pipeline paired with a review by six medical experts. CheXpercept contains 10,400 QA items derived from 2,100 CXRs, covering seven clinically critical pulmonary and cardiac lesions. To demonstrate the current landscape of VLM perception, we benchmark 14 general and medical VLMs on CheXpercept. The models achieve adequate performance only at the coarse level, with accuracy degrading precipitously on deeper visual tasks. Notably, medical VLMs show almost no perceptual advantage over their general-domain counterparts, highlighting a systemic flaw in current domain adaptation. The code and dataset will be publicly available.