Domain-adaptive Zero-Shot Image Enhancement via Locality-Constrained Diffusion Guidance
Authors: Theresa Neubauer, Dimitrios Lenis, Astrid Berg, Maria Wimmer, Gaia Romana De Paolis, Philip Matthias Winter, David Major, Johannes Novotny, +2 more
Denoising Diffusion Probabilistic Models have shown remarkable performance in unconditional image generation. In order to generate images with desired semantics, recent works have restricted the solution space by using guidance constraints in the diffusion sampling process. However, for image enhancement across different domains, these methods struggle to balance two main requirements: looking realistic in the target domain (photorealistic images) and preserving relevant features of the source domain, e.g., low-quality renderings or art paintings. Here, small local changes can alter the fidelity of the image completely, while large changes in other regions might be insignificant. We introduce LocDiff, a locality-constrained guidance method for image enhancement, which serves as a zero-shot extension to pre-trained diffusion models, ensuring the preservation of critical features during domain adaptation. In this way, we retain important local features, while allowing less critical regions to remain unconstrained and not interfere with the guidance process for relevant regions. We evaluate our method on two different domain-shift tasks: For art-to-photo translation, we apply the method in a fully zero-shot setting, preserving facial identity from paintings while generating photorealistic details. For enhancing low-quality fetal ultrasound renderings, we demonstrate zero-shot inference with auxiliary prior alignment. Here, the objective is to artificially add high-resolution characteristics and produce photorealistic ultrasound renderings, a target domain for which no ground truth distribution exists. Our experimental results demonstrate that LocDiff achieves favorable realism-faithfulness trade-offs compared to state-of-the-art methods, enabling controllable cross-domain enhancement.
Figures & tables
Figure 1: Our proposed locality-constrained guidance method named LocDiff controls the generation process of diffusion models with region-adaptive conditions, allowing the user fine-grained control of image enhancement, even under domain shifts. Left: We enhance pre-trained diffusion models with LocDiff to preserve facial characteristics of low-quality fetal ultrasound renderings. Right: Locality-constrained guidance to enhance face parts of art paintings with different strengths (constrained face parts in brackets).
Figure 2: Visualization of the locality-constrained guidance step. For each conditioning region Mk (user-defined or model generated binary masks), the distance function dk measures the distance between the input y and the predicted x0∣t restricted to Mk . For each region Mk , the gradient of dk with respect to xt is computed independently and updates the current xt−1 exclusively within the region Mk .
faithfulness
realism
traditional metrics
Conditioning
LMD ↓
LMD
LMD
LMD
SSIM-F ↑
IDS ↓
FID ↓
LPIPS ↓
PSNR ↑
eye ↓
mouth ↓
nose ↓
WikiArt-Faces Dataset
(I) w/o guidance
8.28
7.78
8.61
6.80
0.41
65.98
56.22
0.36
22.33
(II) mouth & nose
7.15
6.14
4.57
3.64
0.49
55.28
65.97
0.33
24.08
(III) eyes
7.58
3.76
8.13
5.61
0.50
57.20
64.82
0.33
23.68
Table 1: Ablation study on the effectiveness of locality-constrained guidance for different image regions on WikiArt-Faces and MetFaces. Our proposed approach, locality-constrained guidance for multiple regions, demonstrates superior performance in terms of identity preservation, as evidenced by the LMD and IDS metrics. Furthermore, it exhibits an improvement in image quality (FID) compared to full image guidance. The best results are highlighted in bold .
Figure 3: Compared to the full-image guidance (c) , the locality-constrained guidance (b) preserves the natural eye colour and lip structure of the original painting (a) , while the reduced guidance in the background region enhances high-frequency details for hair.
Figure 4: Demonstration of the locality-constrained guidance for face restoration of y . Images x0∣t are generated during the reverse sampling process using no locality-constrained guidance (a) , locality-constrained guidance for eyes only (b) , locality-constrained guidance for eye, eyebrow, mouth and nose, and skin (c) .
Figure 5: Art-to-photo: Qualitative results of our ablation study for single and multi-condition guidance of specific face regions using our locality-constrained guidance. Each color in the attribution map indicates a separate image region that is individually conditioned, where black regions are not guided by our locality-constrained guidance method.
hyperparameter
faithfulness
realism
overall
Tstart
m
d
LMD ↓
SSIM-F ↑
IDS ↓
FID ↓
RF ↓
60
2
d1
2.96 / 3.29
0.61 / 0.63
35.2 / 36.7
72.3 / 81.3
28.2 / 30.0
100
2
d1
2.96 / 3.52
0.61 / 0.62
36.5 / 39.1
67.1 / 72.3
26.5 / 27.4
1 40
2
d 1
3.21 / 3.91
0.60 / 0.61
41.0 / 43.5
64.8 / 70.9
2 6.1 / 27.4
180
2
d1
3.39 / 3.97
0.59 / 0.61
42.0 / 44.8
69.8 / 76.0
28.4 / 29.8
220
2
d1
3.29 / 3.91
0.60 / 0.61
42.5 / 44.9
66.1 / 74.0
26.6 / 28.7
Table 2: Ablation study on the number of diffusion steps Tstart , guidance strength m , and distance function d for the locality-constrained guidance, evaluated on MetFaces/WikiArt datasets. Stronger guidance (lower m indicates stronger guidance) improves faithfulness but harms realism. The best setting is underlined.
Figure 6: Comparison of reconstructed results without locality-constrained guidance across different starting time steps ( Tstart ) illustrating the faithfulness-realism tradeoff. At Tstart=60 , eye appearance lacks realism, while at Tstart=100 , mouth appearance lacks faithfulness. Our proposed locality-constrained guidance method, LocDiff, preserves facial details and maintains realism.
Figure 7: Rendering-to-photo: Qualitative comparison to baseline methods. LocDiff preserves facial characteristics better, as evidenced by the highlighted regions. Zoom in for best view.
faithfulness
realism
traditional metrics
Methods
SSIM-F ↑
IDS ↓
FID ↓
LPIPS ↓
PSNR ↑
DifFace [ 56 ]
0 .73 ± 0.06
5 1.0 ± 7.9
74.8
0 .194 ± 0.04
2 8.27 ± 1.2
PGDiff [ 54 ]
0.69 ± 0.06
54.9 ± 8.1
76.5
0.240 ± 0.05
26.61 ± 1.1
LocDiff (Ours)
0.78 ± 0.05
45.4 ± 7.7
7 5.2
0.188 ± 0.04
28.75 ± 1.2
Table 3: Quantitative comparison on the fetal ultrasound rendering dataset. Best results in bold ; second-best u nderlined.
faithfulness
realism
traditional metrics
Methods
LMD ↓
SSIM-F ↑
IDS ↓
FID ↓
LPIPS ↓
PSNR ↑
WikiArt-Faces dataset
Art2Real [ 49 ]
6.6 ± 8.1
0.50 ± 0.1
51.5 ± 11.3
130.4
0 .27 ± 0.07
18.7 ± 3.1
ILVR [ 7 ]
4.5 ± 5.7
0.56 ± 0.1
46.0 ± 7.1
108.6
0.26 ± 0.06
2 6.7 ± 2.0
DifFace [ 56 ]
5.1 ± 7.2
0.53 ± 0.1
50.8 ± 7.6
66.3
0.30 ± 0.05
25.3 ± 1.8
PGDiff [ 54 ]
4.7 ± 7.0
0.51 ± 0.1
47.7 ± 7.7
6 8.2
0.31 ± 0.06
24.9 ± 2.0
Table 4: Quantitative comparison to SOTA methods on WikiArt-Faces and MetFaces. Best results in bold ; second-best u nderlined.
Methods
Time (s)
Peak GPU Memory (GB)
Art-to-photo task (WikiArt/MetFaces)
Art2Real
0.14 ± 0.01
8.3GB
ILVR
0.64 ± 0.02
2.7GB
DifFace
7.20 ± 0.17
5.8GB
PGDiff
103.47 ± 14.28
4.8GB
DifBIR
15.82 ± 0.59
13.9GB
Table 5: Computational requirements of baseline methods. Time: average inference time per sample.
Figure 8: Art-to-photo: Qualitative comparison with state-of-the-art methods. Our proposed LocDiff enhances the low-quality art painting with photo-realistic details, while preserving facial identity.
Figure 9: Visual comparison with baseline methods. Although DiffBIR excels at denoising and preserving structure, it lacks photorealistic details in the mouth (bottom) and eye (top) regions and retains a painting-like appearance. In contrast, our method (LocDiff) achieves photorealistic domain transfer while maintaining competitive identity preservation.
Table 7: Sensitivity to mask perturbations on MetFaces and Ultrasound datasets. Metrics (mean ± std) show LocDiff is robust to common parsing errors, with performance degrading less than 10% in most cases. Metric changes reflect how perturbations redistribute pixels between regions with different conditioning strengths (MetFaces: strong eyes/skin, weak background; Ultrasound: strong background, weak eyes). Perturbations reducing strongly-conditioned regions trade faithfulness for photorealism, while the reverse occurs when strongly-conditioned regions expand. Visual examples in Figs. 12 and 13 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Impact of parsing disagreement on challenging case. MetFaces sample with lowest cross-model mIoU (0.709) in the test set. Parsing disagreement around the mouth region (mouth vs. skin vs. background) produces visible differences in enhancement: stronger conditioning on mouth/skin regions preserves the moustache better than weak background conditioning. Despite incorrect parsing, no artifacts are introduced in the enhanced images.
Figure 11: Impact of parsing disagreement on challenging ultrasound sample. Ultrasound rendering with lowest cross-model mIoU (0.424) in the test set. FaRL (CelebM-HQ) incorrectly assigns eye and background regions to skin. Given the ultrasound-specific conditioning (strong background, weak skin/eyes), misclassified background regions are slightly more enhanced and the eye region is enhanced differently. Despite these mask errors, enhancement quality degrades moderately with only localized differences.
Figure 12: Mask perturbation robustness analysis on ultrasound data. Top row: input rendering, unperturbed mask, and baseline enhancement. Rows 2-4 show mask perturbation pairs: perturbed mask, resulting enhancement, and RGB difference magnitude from baseline enhancement. Differences are most visible in regions where perturbations alter conditioning strength. For the ultrasound task, background receives strong conditioning while skin/eyes receive weak conditioning. Difference maps reveal that most perturbations cause subtle, localized changes only, with missing eye classes and severe morphological errors producing the most visible deviations.
Figure 13: Mask perturbation robustness analysis on MetFaces. Top row: input, baseline mask, and baseline enhancement. Subsequent rows show pairs of perturbed masks, corresponding enhancements, and difference maps (RGB magnitude from baseline). Strongest differences occur in perturbed regions where conditioning strength changes, particularly when critical facial features like eyes are misclassified, though overall enhancement remains stable. The dilation perturbation better preserves hair structure, as the dilated skin class (more strongly conditioned than the background) partially covers the hair regions.
Figure 14: Failure cases of LocDiff . Top row: original images. Bottom row: enhanced results. The method struggles when: (1) extreme anatomical deviations (elongated ear) are normalized to common shapes present in the training data, (2) heavy artistic brush strokes cause structural misinterpretation (nose deformation), (3) highly abstract face semantics produce only denoising without photorealistic detail generation, and (4) out-of-distribution elements like hands are distorted due to absence from the FFHQ training prior.
Zero-shot domain adaptation is a method for adapting a model to a target domain without utilizing target domain image data. To enable adaptation without target images, existing studies utilize CLIP's embedding space and text description to simulate target-like style features. Despite the previous achievements in zero-shot domain adaptation, we observe that these text-driven methods struggle to capture complex real-world variations and significantly increase adaptation time due to their alignment process. Instead of relying on text descriptions, we explore solutions leveraging image data, which provides diverse and more fine-grained style cues. In this work, we propose SIDA, a novel and efficient zero-shot domain adaptation method leveraging synthetic images. To generate synthetic images, we first create detailed, source-like images and apply image translation to reflect the style of the target domain. We then utilize the style features of these synthetic images as a proxy for the target domain. Based on these features, we introduce Domain Mix and Patch Style Transfer modules, which enable effective modeling of real-world variations. In particular, Domain Mix blends multiple styles to expand the intra-domain representations, and Patch Style Transfer assigns different styles to individual patches. We demonstrate the effectiveness of our method by showing state-of-the-art performance in diverse zero-shot adaptation scenarios, particularly in challenging domains. Moreover, our approach achieves high efficiency by significantly reducing the overall adaptation time.
We study image inpainting with generative diffusion models. Existing methods typically either train dedicated task-specific models, or adapt a pretrained diffusion model separately for each masked image at deployment. We introduce a middle-ground model, termed Amortized Inpainting with Diffusion (AID), which keeps a pretrained diffusion backbone fixed, trains a small reusable guidance module offline, and then reuses it across masked images without per-instance optimization. We formulate it as a deterministic guidance problem with a supervised terminal objective. To make this problem learnable in high dimensions, we derive an auxiliary Gaussian formulation and prove that solving this randomized problem recovers the optimal deterministic guidance field. This bridge yields a principled continuous-time actor--critic algorithm for learning the guidance module in a fully data-driven manner. Empirically, on AFHQv2 and FFHQ under the pixel EDM pipeline and on ImageNet under the latent EDM2 pipeline, AID consistently improves the quality--speed trade-off over strong fixed-backbone and amortized inpainting baselines across multiple mask types, while adding less than one percent trainable overhead.
Yilie Huang, Xun Yu Zhou
Department of Industrial Engineering and Operations Research, Columbia University, New York, NY 10027, USA. · Department of Industrial Engineering and Operations Research & Data Science Institute, Columbia University, New York, NY 10027, USA.
Unconditional diffusion models offer powerful generative priors, yet steering them toward aesthetically enhanced outputs remains largely unexplored. We show that h-space patching, the dominant paradigm for training-free diffusion editing, systematically fails for global, low-level transformations required for aesthetic and perceptual refinement. We introduce a novel, generalized framework for image-editing in unconditional diffusion models without explicit training. This inference-time mechanism operates on low-level features by extracting degradation concept vectors and combining bottleneck patching with classifier-free guidance to guide sampling away from the degraded manifold, producing consistently improved images without any model retraining.