Integrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis
Authors: S. M. Nasif Uddin, Rusab Sarmun, Muhammad E. H. Chowdhury, Adam Mushtak, Israa Al-Hashimi, Sohaib Bassam Zoghoul
Organizations: Department of Electrical and Electronic Engineering, Ahsanullah University of Science and Technology, Bangladesh · Department of Electrical and Electronic Engineering, University of Dhaka, Bangladesh · Department of Electrical Engineering, Qatar University, Doha 2713, Doha, Qatar · Department of Radiology, Hamad Medical Corporation, Doha, Qatar
Primary bone tumors are rare but clinically aggressive neoplasms whose diagnosis from radiographs is challenged by heterogeneous morphology, subtle lesion margins, and overlapping bone structures. To address the limitations of existing single-view models, we present a dual-input, multi-task learning framework that, to our knowledge, is the first to apply bidirectional cross-modal attention between a lesion crop and the full radiograph for joint segmentation and subtype classification. Using the multi-institutional Bone Tumor X-ray Radiograph Dataset (BTXRD, n=3,746), we employ a YOLO-based detector to generate regions of interest, which are paired with full images as inputs to a dual-stream DenseNet121 architecture. Features are integrated via a novel cross-modal attention fusion strategy, refined by Hierarchical Multi-scale Feature Fusion, effectively balancing fine-grained lesion detail with global anatomical context. Evaluated on a held-out patient-level test split, the model demonstrates superior performance over single-input baselines, achieving an overall Dice Similarity Coefficient of 0.896 and a macro-averaged classification F1-score of 0.928. Notably, the system exhibits exceptional sensitivity for malignant osteosarcoma (AUC 0.999), validating the potential of dual-stream context modeling to support radiologists in accurate, early decision-making.
Figures & tables
Figure 1: Class distribution of the BTXRD dataset
Parameter
Value
Model Variant
YOLOv11x
Input Resolution
640×640
Batch Size
8
Training Epochs
118
Optimizer
SGD
Initial Learning Rate
1×10−3
Table 1: YOLOv11 training configuration.
Figure 2: Training-set class distribution before and after augmentation
School of Computer Engineering, Iran University of Science and Technology, Tehran, Iran · School of Computer Engineering, Iran University of Science and technology, Tehran, Iran