cs.CVSep 28, 2026

Style-Driven Data Synthesis and Degradation-Aware Enhancement for Ultrasound Image Restoration

Authors: Yu-Kai Wang, Chun-Xin Tan, Manh-Hung Nguyen, Ching-Chun Huang

Organizations: Department of Computer Science, National Yang Ming Chiao Tung University, Hsinchu, Taiwan · Ho Chi Minh City University of Technology and Engineering, Ho Chi Minh City, Vietnam

Abstract

Low-cost handheld ultrasound devices can be widely deployed compared to professional hospital ultrasound machines. However, their images suffer from compound degradation that can mislead clinical judgment. Motivated by this observation, mapping handheld low-quality (LQ) to hospital high-quality (HQ) images has been considered a valuable research question. Conventionally, the mapping requires pixel-aligned LQ-HQ pairs. This requirement is unsatisfactory in practical scenarios because real scans at different times are never pixel-aligned. This paper addresses the challenge with a two-stage framework. The first stage generates pixel-aligned LQ-HQ datasets, and the second stage trains an enhancement model that improves LQ images. The first stage trains a cycle-consistent style-transfer model on unaligned real LQ-HQ pairs to learn a HQ-to-LQ model. Then, the model transforms real HQ images into pixel-aligned LQ images. Based on the dataset generated by the first stage, the second stage uses the Dual Degradation-Guided (DDG) Low-Rank Adaptation (LoRA) method to fine-tune an LQ-to-HQ model based on aligned pairs. In this stage, the model is based on the well known PiSA-SR framework but inserts a degradation-conditioned correction matrix. Experimental results on the USenhance2023 dataset show that the FID metric is improved by 16.7% over the strongest baseline while other metrics indicate that our enhanced outputs are well aligned with the real HQ distribution. The source code of our method is available at https://github.com/Jason0411202/DDG_LoRA.

Figures & tables

Explore similar work

May 13, 2026cs.CV

HIR-ALIGN: Enhancing Hyperspectral Image Restoration via Diffusion-Based Data Generation

Hyperspectral image (HSI) restoration is crucial for reliable analysis, as real-world HSIs suffer from noise, blur, and resolution loss. However, existing models trained on source data often fail on target domains lacking clean references, a common real-world scenario. To address this, we present HIR-ALIGN, a plug-and-play target-adaptive augmentation framework that enhances HSI restoration by augmenting limited training images with synthetic data matching the target distribution, without extra clean target-domain HSI data. It has three stages: (i) proxy generation, where off-the-shelf restoration models are applied to degraded target observations to produce semantics-preserving proxy HSIs that approximate clean target-domain images; (ii) distribution-adaptive synthesis, where a blur-robust unCLIP diffusion model generates target-aligned RGBs from proxy RGBs with prompt conditioning and embedding-space noise initialization. The warp-based spectral transfer module then synthesizes HSIs by aligning each generated RGB with its proxy RGB, estimating soft patch-wise transport weights, and applying these weights and learnable local interpolation kernels to the proxy HSI; and (iii) aligned supervised finetuning, where restoration networks pretrained on the source distribution are finetuned with proxy HSIs and synthesized target-aligned HSIs, then deployed on degraded target images. We also provide theoretical analysis showing that, under stated assumptions, the proposed augmentation-based finetuning obtains a tighter target-domain restoration-risk upper bound by jointly improving target-distribution coverage and controlling spectral bias. Experiments on simulated and real datasets across denoising, super-resolution, and other restoration tasks demonstrate that HIR-ALIGN is superior to proxy-only target-adaptation baselines and outperforms representative unsupervised methods in most cases.
May 8, 2026eess.IV

A Paired Point-of-Care Ultrasound Dataset for Image Quality Enhancement and Benchmarking via a cGAN Baseline

Purpose: We aim to enhance the image quality of point-of-care ultrasound (POCUS) devices using deep learning and a novel paired dataset of POCUS and high-end ultrasound images. Approach: We collected the first accurately paired dataset using a custom-built automated gantry system of low-end POCUS and high-end ultrasound images. A conditional generative adversarial network (cGAN) was utilized based on the pix2pix architecture, with a U-Net generator that incorporates both L1 and structural similarity index (SSIM) losses to improve perceptual quality. Pretraining on a simulation dataset further boosts performance. Evaluation was performed on 1064 paired ex vivo tissue and phantom ultrasound image sets. Results: Our approach improves the SSIM from 0.29 to 0.54 and PSNR from 19.16 dB to 22.41 dB. No-reference metrics also indicate substantial enhancement, with the Natural Image Quality Evaluator (NIQE) and Perception-based Image Quality Evaluator (PIQE) scores dropping from 7.95 to 4.44 and 31.12 to 19.99, respectively. Conclusions: This work presents the first publicly available accurately paired dataset of low-end POCUS to high end ultrasound images. Additionally, our results demonstrate the potential of the proposed framework to overcome hardware limitations of handheld POCUS, enhancing its diagnostic value in low-resource and point-of-care settings. The POCUS-IQ Dataset is publicly available at https://github.com/NKI-MedTech-AI/POCUS-IQ.
Jun 30, 2026cs.CV

Language-Assisted Super-Resolution from Real-World Low-Resolution Patches

Single image super-resolution aims to reconstruct high-resolution (HR) images from low-resolution (LR) inputs. Training SR models typically requires paired HR-LR data, which is difficult to obtain in reality. As a result, most methods synthesize LR images by artificially degrading HR images with handcrafted kernels or camera ISP adjustments. However, these synthetic degradations fail to capture the complexity of real LR images, leading to poor generalization in practice. To address this, we observe that even within a single high-quality image, regions at different depths exhibit varying resolutions, where distant regions act as LR patches and closer ones as HR patches. This allows the extraction of real, degradation-induced LR patches from real images. Since these LR patches lack paired HR counterparts, we propose LA-SR (Language Assistant for SR), a novel framework for unpaired SR. The key idea of LA-SR is to redefine unpaired SR in the language space, using vision-language models to bridge the LR-HR gap. LA-SR projects images into a semantically rich space representing both content and quality, and applies two language-guided losses: linguistic content loss to preserve semantic fidelity, and linguistic quality loss to enhance perceptual realism. With this alignment, LA-SR effectively super-resolves real LR inputs, producing realistic outputs that overcome the limitations of synthetic-data-trained methods.