cs.CVApr 28, 2026

Robustness of Transformer-Based Fluence Map Prediction Under Clinically Realistic Perturbations

Authors: Ujunwa MgbohRafi Ibn SultanJoshua KimKundan ThindDongxiao Zhu

Abstract

Learning-based fluence map prediction offers a fast alternative to iterative inverse planning in intensity-modulated radiation therapy (IMRT), but its robustness under realistic distribution shifts remains unclear. We study a two-stage transformer pipeline that maps anatomy (CT and contours) to dose and then to beamlet fluence maps. We compare fluence-stage transformer backbones with hierarchical, global, and hybrid attention, trained with a physics-informed loss enforcing energy consistency. Robustness is evaluated under geometric perturbations, radiometric noise, reduced training data, and domain shifts using a prostate IMRT dataset, with additional evaluation of the dose stage on public datasets. Results show smooth degradation under moderate perturbations but sharp failures under severe rotations and noise. Hierarchical transformers (e.g., SwinUNETR) exhibit slower growth in upper-quartile energy error, indicating improved robustness. We further show that SSIM alone fails to capture clinically relevant errors, highlighting the need for physics-informed evaluation.

Explore similar work

May 6, 2026eess.IV

Tumor-aware augmentation with task-guided attention analysis improves rectal cancer segmentation from magnetic resonance images

Although self-supervised pretraining is expected to learn broadly transferable representations, its effectiveness across imaging modalities substantially different from the pretraining domain, and on complex tumor-segmentation tasks, remains understudied. Evaluating CT-pretrained transformers on MRI rectal cancer segmentation, we identified two interacting failure modes in CT-to-MRI transfer: (a) inefficient token usage caused by zero-padding to match pretrained input dimensions, and (b) ineffective feature adaptation. We investigated these vulnerabilities using two primary CT-pretrained hierarchical shifted-window transformer backbones, SMIT and Swin UNETR, together with VoCo as a large-scale-pretrained supporting benchmark; these models differ in pretraining objectives and datasets. Mechanistic analysis leveraged an attention dilution index (ADI), an entropy-based metric quantifying attention diverted toward uninformative padding tokens, and centered kernel alignment (CKA) to measure feature reuse during MRI adaptation. ADI increased with zero-padding, while high feature reuse did not necessarily translate to improved downstream accuracy. To mitigate these issues, we introduced two interventions: a tumor-aware augmentation strategy to expand tumor appearance heterogeneity coverage, and an anisotropic cropping strategy to restore token efficiency. Fine-tuning with these strategies on identical rectal MRI datasets yielded detection rates of 91.1% (225/247) and 88.7% (219/247) for the primary SMIT and Swin UNETR backbones, with the supporting VoCo benchmark reaching 90.3% (223/247), demonstrating significantly improved robustness under CT-to-MRI transfer. This study is among the first to examine when pretrained transformers fail to transfer across imaging modalities and demonstrates how targeted mitigation strategies can systematically overcome cross-modality transfer limitations.
Aneesh Rangnekar, Joao Miranda, Natally Horvat +13
Aug 10, 2026cs.CV

DoseBridge: Denoising Diffusion Bridge Model for Dose Prediction in Lung Intensity-Modulated Proton Therapy

Most radiotherapy dose-prediction models use only CT images and anatomical structures, although intensity-modulated proton therapy (IMPT) dose also depends strongly on beam geometry and available clinical datasets are often small. We present DoseBridge, a denoising diffusion bridge model that uses the patient CT as a structured bridge endpoint and encodes plan-specific beam geometry in a spatially aligned beam mask. Multiscale fusion combines CT, target, organ-at-risk, and beam-mask representations with 1.95% additional parameters. DoseBridge was retrospectively evaluated on single-institution CT images and treatment plans from 52 patients with advanced-stage lung cancer treated with 60 Gy in 30 fractions; 42 cases were used for training and 10 for testing. Performance was assessed using image-similarity, dose-volume, and Lyman-Kutcher-Burman normal-tissue complication probability (NTCP) metrics and compared with two deep-learning models. On the test cohort, DoseBridge achieved a mean absolute error of 4.170 Gy, peak signal-to-noise ratio of 23.06 dB, and structural similarity index of 0.798, outperforming both comparison models on these metrics. Clinical target volume D95 differed from the reference dose by 0.62 +/- 1.6 Gy; signed organ-at-risk mean-dose differences ranged from -0.32 to 0.24 Gy, and NTCP differences were -0.40 +/- 2.2 and 0.52 +/- 3.4 percentage points for acute esophagitis and radiation pneumonitis, respectively. Changing only the beam mask redirected predicted low-dose entrance regions while preserving the high-dose target region. To our knowledge, DoseBridge is the first denoising diffusion bridge model for radiotherapy dose prediction. These results support its feasibility as a beam-aware planning prior for lung IMPT, pending evaluation in larger external cohorts.
Zerun Zhang, Xiaoda Cong, Xiangkun Xu +2
May 14, 2026cs.CV

Towards Real-Time Autonomous Navigation: Transformer-Based Catheter Tip Tracking in Fluoroscopy

Purpose: Mechanical thrombectomy (MT) improves stroke outcomes, but is limited by a lack of local treatment access. Widespread distribution of reinforcement learning (RL)-based robotic systems can be used to alleviate this challenge through autonomous navigation, but current RL methods require live device tip coordinate tracking to function. This paper aims to develop and evaluate a real-time catheter tip tracking pipeline under fluoroscopy, addressing challenges such as low contrast, noise, and device occlusion. Methods: A multi-threaded pipeline was designed, incorporating frame reading, preprocessing, inference, and post-processing. Deep learning segmentation models, including U-Net, U-Net+Transformer, and SegFormer, were trained and benchmarked using two-class and three-class formulations. Post-processing involved two-step component filtering, one-pixel medial skeletonization, and greedy arc-length path following with contour fall-back. Results: On manually-labeled moderate complexity fluoroscopic video data, the two-class SegFormer achieved a mean absolute error of 4.44 mm, outperforming U-Net (4.60 mm), U-Net+Transformer (6.20 mm) and all three-class models (5.19-7.74 mm). On segmentation benchmarks, the system exceeded state-of-the-art CathAction results with improvements of up to +5% in Dice scores for three-segmentation. Conclusion: The results demonstrate that the proposed multi-threaded tracking framework maintains stable performance under challenging imaging conditions, outperforming prior benchmarks, while providing a reliable and efficient foundation for RL-based autonomous MT navigation.
Harry Robertshaw, Yanghe Hao, Weiyuan Deng +6