Organizations: Department of Computer Science, National Yang Ming Chiao Tung University, Hsinchu, Taiwan · Ho Chi Minh City University of Technology and Engineering, Ho Chi Minh City, Vietnam
Low-cost handheld ultrasound devices can be widely deployed compared to professional hospital ultrasound machines. However, their images suffer from compound degradation that can mislead clinical judgment. Motivated by this observation, mapping handheld low-quality (LQ) to hospital high-quality (HQ) images has been considered a valuable research question. Conventionally, the mapping requires pixel-aligned LQ-HQ pairs. This requirement is unsatisfactory in practical scenarios because real scans at different times are never pixel-aligned. This paper addresses the challenge with a two-stage framework. The first stage generates pixel-aligned LQ-HQ datasets, and the second stage trains an enhancement model that improves LQ images. The first stage trains a cycle-consistent style-transfer model on unaligned real LQ-HQ pairs to learn a HQ-to-LQ model. Then, the model transforms real HQ images into pixel-aligned LQ images. Based on the dataset generated by the first stage, the second stage uses the Dual Degradation-Guided (DDG) Low-Rank Adaptation (LoRA) method to fine-tune an LQ-to-HQ model based on aligned pairs. In this stage, the model is based on the well known PiSA-SR framework but inserts a degradation-conditioned correction matrix. Experimental results on the USenhance2023 dataset show that the FID metric is improved by 16.7% over the strongest baseline while other metrics indicate that our enhanced outputs are well aligned with the real HQ distribution. The source code of our method is available at https://github.com/Jason0411202/DDG_LoRA.
Figures & tables
Figure 1 : Style-driven data synthesis pipeline. Stage 1 splits USenhance2023 into a 4:1 train/test partition (1), trains a style transfer model G on unaligned real LQ–HQ pairs (2), applies G to each real HQ image yi to obtain a pixel-aligned LQ counterpart x~i (3), and assembles the USenhance2023-Aligned dataset (4). Stage 2 trains the DDG-LoRA enhancer on this dataset (5).
Figure 2 : DDG-LoRA architecture. The pixel stage trains Δθpix , then the semantic stage freezes it and trains Δθsem . Within each stage, DDG-LoRA replaces the static LoRA update with a degradation-conditioned variant driven by a descriptor d that the DE network extracts from the LQ image xL . The semantic stage further drops the CSD loss LCSD .
Training LQ Source
FID ↓
NIQE ↓
PI ↓
Tenengrad ↑
Entropy ↑
realesrgan_deg [ 14 ]
127.01
6.038
5.598
0.0608
6.233
physics_guided_deg [ 8 ]
127.17
5.485
5.198
0.0618
6.320
usbsr_deg [ 7 ]
156.07
6.824
6.635
0.0405
6.287
ultrasound_deg
185.75
7.909
7.377
0.0336
6.163
USenhance2023-Aligned (ours)
99.00
5.619
4.764
0.1043
6.861
Table 1 : Comparison of DDG-LoRA (ours) trained with different LQ data sources, evaluated on the real-world USenhance2023 test set. Red = best, Blue = second best.
Figure 3 : Visual comparison of all methods in Table 2 on a real-world USenhance2023 sample. The LQ input is shown in (a), the enhanced output of each method in (b)–(i), and the GT reference in (j). Note that because both the LQ and the GT are real-world scans, they are not pixel-aligned.
Method
FID ↓
NIQE ↓
PI ↓
Tenengrad ↑
Entropy ↑
EDSR [ 10 ]
187.57
9.058
7.541
0.0573
6.360
CycleSR [ 2 ]
192.36
8.293
7.139
0.0489
6.512
DARSR [ 19 ]
235.03
8.883
7.466
0.0680
6.341
PDM [ 11 ]
183.92
8.236
7.416
0.0317
6.415
USBSR [ 7 ]
168.33
8.416
7.609
0.0426
6.530
OSEDiff [ 16 ]
146.80
6.305
5.837
0.0531
6.273
Table 2 : Overall comparison of different enhancement methods on the real-world USenhance2023 test set. Red = best, Blue = second best.
Configuration
FID ↓
NIQE ↓
PI ↓
Tenengrad ↑
Entropy ↑
w/o DDG-LoRA
98.98
5.589
4.815
0.1006
6.808
w/ DDG-LoRA (ours)
99.00
5.619
4.764
0.1043
6.861
Table 3 : Ablation study on the degradation conditioning module in DDG-LoRA. Bold red indicates the better result for each metric.
Hyperspectral image (HSI) restoration is crucial for reliable analysis, as real-world HSIs suffer from noise, blur, and resolution loss. However, existing models trained on source data often fail on target domains lacking clean references, a common real-world scenario. To address this, we present HIR-ALIGN, a plug-and-play target-adaptive augmentation framework that enhances HSI restoration by augmenting limited training images with synthetic data matching the target distribution, without extra clean target-domain HSI data. It has three stages: (i) proxy generation, where off-the-shelf restoration models are applied to degraded target observations to produce semantics-preserving proxy HSIs that approximate clean target-domain images; (ii) distribution-adaptive synthesis, where a blur-robust unCLIP diffusion model generates target-aligned RGBs from proxy RGBs with prompt conditioning and embedding-space noise initialization. The warp-based spectral transfer module then synthesizes HSIs by aligning each generated RGB with its proxy RGB, estimating soft patch-wise transport weights, and applying these weights and learnable local interpolation kernels to the proxy HSI; and (iii) aligned supervised finetuning, where restoration networks pretrained on the source distribution are finetuned with proxy HSIs and synthesized target-aligned HSIs, then deployed on degraded target images. We also provide theoretical analysis showing that, under stated assumptions, the proposed augmentation-based finetuning obtains a tighter target-domain restoration-risk upper bound by jointly improving target-distribution coverage and controlling spectral bias. Experiments on simulated and real datasets across denoising, super-resolution, and other restoration tasks demonstrate that HIR-ALIGN is superior to proxy-only target-adaptation baselines and outperforms representative unsupervised methods in most cases.
Purpose: We aim to enhance the image quality of point-of-care ultrasound (POCUS) devices using deep learning and a novel paired dataset of POCUS and high-end ultrasound images. Approach: We collected the first accurately paired dataset using a custom-built automated gantry system of low-end POCUS and high-end ultrasound images. A conditional generative adversarial network (cGAN) was utilized based on the pix2pix architecture, with a U-Net generator that incorporates both L1 and structural similarity index (SSIM) losses to improve perceptual quality. Pretraining on a simulation dataset further boosts performance. Evaluation was performed on 1064 paired ex vivo tissue and phantom ultrasound image sets. Results: Our approach improves the SSIM from 0.29 to 0.54 and PSNR from 19.16 dB to 22.41 dB. No-reference metrics also indicate substantial enhancement, with the Natural Image Quality Evaluator (NIQE) and Perception-based Image Quality Evaluator (PIQE) scores dropping from 7.95 to 4.44 and 31.12 to 19.99, respectively. Conclusions: This work presents the first publicly available accurately paired dataset of low-end POCUS to high end ultrasound images. Additionally, our results demonstrate the potential of the proposed framework to overcome hardware limitations of handheld POCUS, enhancing its diagnostic value in low-resource and point-of-care settings. The POCUS-IQ Dataset is publicly available at https://github.com/NKI-MedTech-AI/POCUS-IQ.
Lennard M. van Karnenbeek, Hilde G. A. van der Pol, Mark Wijkhuizen +5
Department of Nanobiophysics, Faculty of Science and Technology, University of Twente, Drienerlolaan 5, 7522 NB Enschede, The Netherlands · Image-Guided Surgery, Department of Surgery, Netherlands Cancer Institute, Plesmanlaan 121, 1066 CX Amsterdam, The Netherlands · Department of Radiology, Netherlands Cancer Institute, Amsterdam, The Netherlands +1
Single image super-resolution aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs. Training SR models typically requires paired HR-LR data, which is difficult to obtain in reality. As a result, most methods synthesize LR images by artificially degrading HR images with handcrafted kernels or camera ISP adjustments. However, these synthetic degradations fail to capture the complexity of real LR images, leading to poor generalization in practice. To address this, we observe that even within a single high-quality image, regions at different depths exhibit varying resolutions, where distant regions act as LR patches and closer ones as HR patches. This allows the extraction of real, degradation-induced LR patches from real images. Since these LR patches lack paired HR counterparts, we propose LA-SR (Language Assistant for SR), a novel framework for unpaired SR. The key idea of LA-SR is to redefine unpaired SR in the language space, using vision-language models to bridge the LR-HR gap. LA-SR projects images into a semantically rich space representing both content and quality, and applies two language-guided losses: linguistic content loss to preserve semantic fidelity, and linguistic quality loss to enhance perceptual realism. With this alignment, LA-SR effectively super-resolves real LR inputs, producing realistic outputs that overcome the limitations of synthetic-data-trained methods.
Joonkyu Park, Kyoung Mu Lee
Dept. of ECE&ASRI, 2IPAI, Seoul National University