A Unified Deep Learning Framework for Motion Correction in Medical Imaging
Authors: Jian Wang, Razieh Faghihpirayesh, Danny Joca, Polina Golland, Ali Gholipour
Organizations: Boston Children’s Hospital and Harvard Medical School, Boston, MA, USA · Department of Electrical and Computer Engineering at Northeastern University and also with the Department of Radiology at Boston Children’s Hospital, Boston, MA, USA · Department of Radiology at Boston Children’s Hospital and Harvard Medical School, Boston, MA, USA · CSAIL (Computer Science and Artificial Intelligence Laboratory) at Massachusetts Institute of Technology (MIT), Cambridge, MA, USA · Departments of Radiological Sciences, Electrical Engineering and Computer Science, and Biomedical Engineering at the University of California Irvine, CA, USA
Deep learning has shown significant value in medical image registration for motion correction; however, current techniques are either limited by the type and range of motion they can handle or require iterative inference and/or retraining for new imaging data. To address these limitations, we introduce UniMo, a Unified Motion Correction framework that uses deep neural networks to correct various types of motion in medical imaging. UniMo uses an alternating optimization scheme with a unified loss function to train an integrated model of 1) an equivariant neural network for global rigid motion correction and 2) an encoder-decoder network for local deformations. It features a geometric deformation augmenter that 1) enhances the robustness of global motion correction by addressing local deformations, whether caused by non-rigid motion or geometric distortions, and 2) generates augmented data to improve training. As a hybrid model that uses both image intensities and shapes, UniMo is robust to appearance variations and generalizes to various imaging modalities without retraining. We trained and tested UniMo for motion tracking in fetal magnetic resonance imaging, which is challenging due to 1) both large rigid and non-rigid motion and 2) large variations in image appearance. We then tested the trained model, without retraining, on three public datasets: MedMNIST, lung CT, and BraTS. UniMo surpassed existing motion correction methods in accuracy and, notably, enabled one-time training on a single modality while maintaining high stability and adaptability across multiple unseen imaging datasets. By offering a unified solution to motion correction, UniMo marks a significant advance in challenging applications with a mixture of bulk motion and local deformations. Code is available at https://github.com/IntelligentImaging/UNIMO
Figures & tables
Fig. 1: An illustration of the network architecture of our proposed motion correction learning framework, UniMo. Top : Motion correction is based on equivariant neural networks that take both images and segmentations as input. The low-dimensional equivariant spatial means of the images and segmentations are estimated simultaneously. An enhanced hybrid rigid transformation Q is computed from both domains. The rigid loss function is calculated between the aligned and target images/labels. Bottom : A geometric shape augmenter (i.e., the deformation correction network described in Section III-C) is incorporated into the U-Net based neural networks. It takes both aligned and target images/segmentations and estimates the transformation fields for both domains using separate loss functions for segmentations and images.
Fig. 2: Convergence of the weight parameter λ during training of UniMo. a–d , Trajectories of the effective weight λt defined by Eq. ( 14 ) over 1000 epochs of training with the loss function ( 11 ), under four initializations λ0=0.15 , 0.05 , 0.50 , and 0.75 (corresponding to α0=3 , 1 , 10 , and 15 with λmax=20 ), respectively. Solid curves show the smoothed trajectory, shaded bands the per-epoch fluctuation envelope, and dashed lines the final selected value λ∗ . Insets: the deviation ∣αt−α∗∣ of the unconstrained variable α=λmaxλ on a logarithmic scale; the deviation of λt is this quantity divided by λmax and follows the same curve. The approximately linear decay over the first ∼500 epochs indicates that λt converges to λ∗ at an exponential rate, after which the residual fluctuation is dominated by stochastic-gradient noise. Convergence is observed regardless of the initialization, whether λ0 is set below or above λ∗ .
Fig. 3: Two examples of heat maps of tSNR estimated from all motion correction models over 50 fetal EPI pairs. From left to right: Target image, heat maps from UniMo, Equiv-Filter, KeyMorph and DeepPose. Higher tSNR values indicate the best alignment of the image time series was obtained from UniMo.
Fig. 4: Statistical results for both translational and angular errors of all models on single modality. Left: motion correction performance reported over 300 image pairs; Right: motion tracking performance on 70 sequences of fetal EPI scans with simulated motions. Boxes show the median and interquartile range of the per-pair (left) and per-sequence (right) errors, whiskers extend to 1.5 times the interquartile range, and circles mark outliers. p -values of the paired comparisons between UniMo and each baseline are given in Table I .
Fig. 5: Statistical results of Dice comparison on fetal EPIs with unknown motions. Left: Motion correction performance with different degrees of motions, small ( Tmax=10mm , Rmax=5∘ ), medium ( Tmax=20mm,Rmax=10∘ ) and large motions ( Tmax=30mm , Rmax=20∘ ). The mean dice score of our best model, for motion levels from left to right are, 0.97, 0.93, 0.92 (violins show the per-pair distribution of Dice values, with the inner marker at the median; p -values against each baseline in Table I ) . Right: motion tracking performance across varying degrees and lengths of image sequences ( T ). Small ( Tmax=10mm , Rmax=5∘ ) and large motions ( Tmax=30mm , Rmax=20∘ ) were evaluated. Report efficiency with average time consumption: 0.501s per pair / 9.960s per sequence when T=20 .
Metric
Equiv-Filter
KeyMorph
DeepPose
Fetal EPI, motion correction ( n=300 pairs)
Translational error
0.003
<0.001
<0.001
Angular error
0.011
0.010
<0.001
Fetal EPI, motion tracking ( n=70 sequences)
Translational error
0.001
0.002
<0.001
Angular error
0.023
0.001
<0.001
TABLE I: p -values of the paired comparison between UniMo and Equiv-Filter [ 75 ] , KeyMorph [ 72 ] , and DeepPose [ 68 ] (two-sided paired Wilcoxon signed-rank test). Values below 0.001 are reported as <0.001 . All comparisons are significant at the 0.05 level and remain so after Holm-Bonferroni correction over the three baselines within each row.
Fig. 6: Motion correction comparison across multiple image modalities for all models. The image modalities from top to bottom are: T1 MRI scans of brains containing lesions, lung CT scans, EPI scans of fetal brains, and the shapes of adrenal glands. The images from left to right are: source, target, and aligned images by UniMo, Equivariant Filter [ 75 ] , KeyMorph [ 72 ] , and DeepPose [ 68 ] . The aligned images are displayed with red contours for better visualization.
Fig. 7: From left to right: translational and angular errors of motion correction of four image modalities, epoch number of best model performance from training with varying training dataset size, and quantitative report of Dice accuracy with varying training dataset size. In the two left panels, boxes show the median and interquartile range of the per-pair errors, whiskers extend to 1.5 times the interquartile range, and circles mark outliers.
Motion artifacts in magnetic resonance imaging (MRI) degrade diagnostic reliability. Existing deep learning methods are typically contrast-specific and fail to generalize across diverse modalities and artifact severities. We propose a unified framework combining parameter-informed contrast disentanglement with severity-aware adaptive correction. ScanCLIP, pretrained on over 30,000 MRI text-image pairs, derives contrast embeddings from acquisition parameters to disentangle contrast style from anatomical content, yielding contrast-free features. A Vision Transformer then estimates motion severity and routes features through a Mixture-of-Experts network, enabling targeted artifact correction. A dual-pathway decoder reconstructs both the clean image and residual artifact map, enforcing image-space consistency. On IXI and HCP benchmarks, our method improves PSNR by 0.75 dB and SSIM by up to 0.0279 over state-of-the-art approaches, with larger gains at higher artifact severities. It further demonstrates robust zero-shot generalization on real-world clinical data acquired with unseen scanning parameters, where existing methods either fail to remove artifacts or introduce additional distortions.
Learning-based medical image registration has matched the accuracy of conventional methods while offering superior computational efficiency. However, existing approaches suffer from poor generalization across diverse clinical scenarios, requiring the laborious development of multiple isolated networks for specific registration tasks, \emph{e.g.}, inter-/intra-subject registration or anatomical region-specific alignment, leading to cumbersome development pipelines. To overcome this limitation, we propose \textbf{UniReg}, the first conditional unified model for multi-scenario medical image registration, which combines the precision advantages of task-specific learning methods with the generalization of traditional optimization methods. Our key innovation is a unified registration framework that adaptively estimates deformation fields conditioned on: (1) anatomical structure priors, (2) registration type constraints (inter/intra-subject), and (3) instance-specific features, enabling effective alignment across heterogeneous CT and MR registration scenarios within a single model. Through comprehensive experiments on multiple CT/MR registration datasets, UniReg achieves superior average registration accuracy compared with current state-of-the-art learning-based methods while exhibiting strong cross-scenario generalization. Moreover, by replacing multiple isolated task-specific models with a compact unified model, UniReg substantially reduces the overall training burden in terms of total training cost and model redundancy.
Zi Li, Jianpeng Zhang, Tai Ma +7
Department of Electrical and Computer Engineering, The University of Hong Kong, Hong Kong SAR, China · DAMO Academy, Alibaba Group · The First Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, China +1
Quantitative cardiac magnetic resonance imaging (MRI) enables non-invasive myocardial tissue characterization but relies on robust motion correction within these variable-length, variable-contrast image sequences. Groupwise registration, which simultaneously aligns all images, has shown greater robustness than pairwise registration for motion correction. However, current deep-learning-based groupwise registration methods cannot generalize across MRI sequences: the architecture typically encodes input data as a fixed-length channel stack, which rigidly couples network design to protocol-specific sequence length, input ordering, and contrast dynamics. At inference time, any change in imaging protocols will render the network unusable. In this work, we introduce \emph{\AnyTwoReg}, a new set-based groupwise registration framework that takes a quantitative MRI sequence as an unordered set. This set formulation fundamentally decouples network design from sequence length and input ordering. By utilizing a shared encoder and correlation-guided feature aggregation, \emph{\AnyTwoReg} constructs a permutation-invariant canonical reference for registration, and learns a permutation-equivariant mapping from images to deformation fields. Additionally, we extract contrast-insensitive image features from an existing foundation model to handle extreme contrast variations. Trained exclusively on a single public T1 mapping dataset (STONE, sequence length L=11), \AnyTwoReg generalizes to two unseen quantitative MRI datasets (MOLLI, ASL) with variable lengths (L∈[11,60]) and different contrast dynamics. It achieves strong cross-protocol generalization in a zero-shot manner, and consistently improves downstream quantitative mapping quality. Notably, while designed for quantitative MRI sequences, our framework is directly applicable to Cine MRI sequences for inter-cardiac-phase registration.
Yi Zhang, Yidong Zhao, Tijmen Toxopeus +3
Department of Imaging Physics, Delft University of Technology, The Netherlands