Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
Organizations: School of Computer Science and Information Engineering, Hefei University of Technology, China
Abstract
Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these Dual image references within Frequency-aware Diffusion Models, dubbed Dual-FDM, to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized image.Extensive experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from https://github.com/htyjers/Dual-FDM.
Figures & tables
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| ID | Both foreground and background | Foreground | Background |
| 1 | A snow leopard is leaping across a fractured frozen expanse. | a snow leopard | the fractured frozen expanse |
| 2 | A giant squid is drifting in a radiant deep chasm. | a giant squid | the radiant deep chasm |
| 3 | A phoenix is glowing over an age-worn conflict field. | the phoenix | the age-worn conflict field |
| 4 | A red-crowned crane is gliding through a crimson-lit horizon. | a red-crowned crane | the crimson horizon |
| 5 | A polar bear is sitting beside a fractured icy dome-like form. | a polar bear | the fractured icy dome |
| 6 | A chameleon is blending into a glare-soaked luminous corridor. | a chameleon | the luminous corridor |
| Method | Foreground | Background | ||||
| CLIP-T | CLIP-I | DINO-I | CLIP-T | CLIP-I | DINO-I | |
| 0.3115 | 0.8228 | 0.5774 | 0.2303 | 0.6157 | 0.1353 | |
| 0.2629 | 0.7090 | 0.2591 | 0.2486 | 0.6421 | 0.1966 | |
| Method | CLIP-T-Mean | CLIP-T-Std | CLIP-I-Mean | CLIP-I-Std |
| 0.2469 (-0.68%) | 0.0352 (+13.5%) | 0.6450 (+0.45%) | 0.0867 (+8.36%) | |
| 0.2496 (+0.40%) | 0.0344 (+11.0%) | 0.6433 (+0.19%) | 0.0806 (+0.62%) | |
| 0.2500 (+0.56%) | 0.0332 (+7.10%) | 0.6401 (-0.31%) | 0.0842 (+5.11%) | |
| 0.2495 (+0.36%) | 0.0334 (+7.74%) | 0.6382 (-0.30%) | 0.0826 (+3.11%) | |
| 0.2499 (+0.52%) | 0.0322 (+3.87%) | 0.6416 (-0.08%) | 0.0868 (+8.36%) | |
| 0.2498 (+0.49%) | 0.0330 (+6.45%) | 0.6401 (-0.31%) | 0.0883 (+10.2%) |