Personalized Text-to-Image (PT2I) generation aims to produce customized images based on reference images. A prominent interest pertains to the integration of an image prompt adapter to facilitate zero-shot PT2I without test-time fine-tuning. However, current methods grapple with three fundamental challenges: 1. the elusive equilibrium between Concept Preservation (CP) and Prompt Following (PF), 2. the difficulty in retaining fine-grained concept details in reference images, and 3. the restricted scalability to extend to multi-subject personalization. To tackle these challenges, we present Dynamic Image Prompt Adapter (DynaIP), a cutting-edge plugin to enhance the fine-grained concept fidelity, CP-PF balance, and subject scalability of state-of-the-art T2I multimodal diffusion transformers (MM-DiT) for PT2I generation. Our key finding is that MM-DiT inherently exhibit decoupling learning behavior when injecting reference image features into its dual branches via cross attentions. Based on this, we design an innovative Dynamic Decoupling Strategy that removes the interference of concept-agnostic information during inference, significantly enhancing the CP-PF balance and further bolstering the scalability of multi-subject compositions. Moreover, we identify the visual encoder as a key factor affecting fine-grained CP and reveal that the hierarchical features of commonly used CLIP can capture visual information at diverse granularity levels. Therefore, we introduce a novel Hierarchical Mixture-of-Experts Feature Fusion Module to fully leverage the hierarchical features of CLIP, remarkably elevating the fine-grained concept fidelity while also providing flexible control of visual granularity. Extensive experiments across single- and multi-subject PT2I tasks verify that our DynaIP outperforms existing approaches, while requiring only single-subject training datasets.
Figures & tables
Figure 1 : Representative results showcase the capabilities of DynaIP in: (a) Scalable zero-shot personalized text-to-image generation —spanning single-subject to multi-subject— trained solely on single-subject datasets . (b) Flexible control on the visual granularity of concept preservation , enabled by modulating fusion coefficients for image features across hierarchical levels. (c) Native compatibility with base model extensions , unlocking diverse application scenarios.
Figure 2 : Limitations of existing adapter-based PT2I methods ( e.g . , [ 72 , 56 , 16 ] ), including (a) irreconcilable trade-off between CP and PF, (b) loss of fine-grained concept details, and (c) restricted scalability to directly extend SS-PT2I to MS-PT2I via mask-guided feature injection. Our proposed DynaIP addresses all these challenges.
Figure 3 : Training/Inference pipeline of (a, c) vanilla IP-Adapter and (b, c) our DynaIP.
Figure 4 : Illustration of the decoupling learning behavior of MM-DiT .
Figure 5 : Left : Architecture of our proposed HMoE-FFM. Right : Personalization results generated by injecting features from different layers of CLIP via cross-attentions, demonstrating that CLIP’s hierarchical features can capture visual information at diverse granularity levels.
Figure 6 : Qualitative comparisons on single- and multi-subject PT2I generation.
Method
Single-subject
Multi-subject
CP
PF
CP ⋅ PF
CP
PF
CP ⋅ PF
DreamBooth
0.458
0.721
0.330
-
-
-
DreamBooth LoRA
0.594
0.840
0.499
-
-
-
Textual Inversion
0.348
0.633
0.220
-
-
-
IP-Adapter-Plus
0.738
0.668
0.493
-
-
-
DisEnvisioner
0.559
0.664
0.371
-
-
-
Table 1 : Quantitative comparisons with SOTA methods. CP: Concept Preservation, PF: Prompt Following. The highest and second-highest scores are highlighted.
Figure 7 : Qualitative ablation study results of Left : Dynamic Decoupling Strategy (DDS) and Right : different feature fusion approaches and token-concat baseline.
Setting
Single-subject
Multi-subject
CP
PF
CP ⋅ PF
CP
PF
CP ⋅ PF
(1) Full Model
0.696
0.934
0.650
0.617
0.997
0.615
(2) w/o DDS
0.785
0.799
0.627
0.499
0.545
0.272
(3) Add Fusion
0.691
0.916
0.633
0.609
0.995
0.606
(4) Concat Fusion
0.692
0.909
0.629
0.607
0.992
0.602
(5) Only Shallow
0.627
0.924
0.579
0.464
0.991
0.460
Table 2 : Quantitative ablation study results.
Figure 8 : Evaluation prompt of MLLMs to evaluate concept-specific and concept-agnostic information.
Type
Prompt
[P]
living+living living+object object+object
a {0} and a {1} [P], {0} on the left, and {1} on the right
in a room
in the snow
in the jungle
on the beach
on the grass
on a cobblestone street
Table 3 : Prompt details of our multi-subject DynaIP-Bench. Each combination type has preset prompts. [P] denotes prompt variations about the scene or actions.
Figure 9 : MLLM evaluation prompts for Concept Preservation (CP) on Left : single-subject and Right : multi-subject PT2I generation tasks.
Figure 10 : MLLM evaluation prompts for Left : Concept Preservation (CP) on single-style PT2I generation tasks and for Right : Prompt Following (PF) on both single- and multi-subject PT2I generation tasks.
Figure 11 : Additional qualitative comparisons on multi-subject PT2I generation.
Figure 12 : Additional qualitative comparisons on single-subject PT2I generation.
Figure 13 : Additional qualitative comparisons on single-subject PT2I generation.
Figure 14 : Exemplar results of prompts with complex interactions of multiple subjects.
Figure 15 : Control on the visual granularity of concept preservation , enabled by modulating fusion coefficients ( [wLow,wMid,wHigh] in Eq. (7) in the main paper) of experts’ outputs in HMoE-FFM. Please zoom in to observe the details.
Figure 16 : Control on the visual granularity of concept preservation , enabled by modulating fusion coefficients ( [wLow,wMid,wHigh] in Eq. (7) in the main paper) of experts’ outputs in HMoE-FFM. Please zoom in to observe the details.
Method
CP
PF
OS
DreamBooth
0.288
0.286
0.171
DreamBooth LoRA
0.709
0.782
0.660
Textual Inversion
0.142
0.198
0.100
IP-Adapter-Plus
0.745
0.442
0.297
DisEnvisioner
0.517
0.437
0.227
FLUX.1 IP-Adapter
0.728
0.496
0.336
Table 4 : A/B test user study results. We report the adjusted advantage ratios ( lose+tiewin+tie ) of the competing methods relative to our method. Lower values indicate that our method performs better relative to the competing method. CP: Concept Preservation, PF: Prompt Following. OS: Overall Satisfaction. The highest and second-highest scores are highlighted.
Figure 17 : More analytical results of HMoE-FFM . Left: We present the fusion coefficients ( [wLow,wMid,wHigh] in Eq. (7) of the main paper), adaptively predicted by the routing module of HMoE-FFM, across diverse input reference images. The highest and second-highest coefficients are highlighted. Right: we visualize how our model adapts to prompt adjustments across diverse style and content variations.
Figure 18 : Personalization results generated by injecting features from more layers of CLIP via cross-attentions.
Figure 19 : Limitation of DynaIP . The performance of concept preservation may degrade if the subjects generated by the base model deviate significantly from the reference subjects.
Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant challenges when handling complex prompt, particularly those involving multiple subjects with distinct attributes. Inspired by the human drawing process, which first outlines the composition and then incrementally adds details, we propose Detail++, a training-free framework that introduces a novel Progressive Detail Injection (PDI) strategy to address this limitation. Specifically, we decompose a complex prompt into a sequence of simplified sub-prompts, guiding the generation process in stages. This staged generation leverages the inherent layout-controlling capacity of self-attention to first ensure global composition, followed by precise refinement. To achieve accurate binding between attributes and corresponding subjects, we exploit cross-attention mechanisms and further introduce a Centroid Alignment Loss at test time to reduce binding noise and enhance attribute consistency. Extensive experiments on T2I-CompBench and a newly constructed style composition benchmark demonstrate that Detail++ significantly outperforms existing methods, particularly in scenarios involving multiple objects and complex stylistic conditions.
Lifeng Chen, Jiner Wang, Zihao Pan +3
AGI Lab, Westlake University · Nanyang Technological University
Subject-driven personalized text-to-image generation requires a pretrained diffusion model to acquire a specific subject from a few reference images while preserving subject identity, following novel text prompts, and maintaining sample diversity. Existing optimization-based methods instantiate subject adaptation through full fine-tuning, textual embedding optimization, or low-rank parameter updates; PaRa further constrains personalization from the perspective of parameter rank reduction. However, a uniform low-rank constraint or a uniform adapter strength cannot explicitly distinguish the capacity requirements of different denoising stages. Moreover, inference-time candidate selection driven mainly by identity similarity may compress the selected samples in the visual representation space. We decompose the problem into two complementary components: SPaRa denotes training-side stage-aware low-rank adaptation, DCAL denotes inference-side distribution-calibrated candidate selection, and SPaRa-DCAL denotes the combined framework. Theoretical analysis shows that timestep-dependent scaling controls the effective perturbation magnitude of a low-rank adapter, while identity-biased candidate selection restricts the radius of selected features around the reference center under explicit conditions. Auditable experiments under the SDXL and DreamBooth 30-subject protocol show that DCAL improves 1-LPIPS, CLIP-I, DINO-I, and CLIP-T on a fixed LoRA candidate pool, while revealing a clear trade-off with CLIP/DINO pairwise diversity and pairwise LPIPS. These results indicate that personalized generation should be evaluated through identity consistency, text alignment, and representation diversity rather than identity metrics alone.
Wenyan Xu, Alizer Wong
School of Computer Science, Guangdong University of Technology · School of Computer Science, Peking University · 3ManXis
Multi-subject image generation aims to synthesize images that faithfully preserve the identities of multiple reference subjects while following textual instructions. However, existing methods often suffer from identity inconsistency and limited compositional control, as they rely on diffusion models to implicitly associate text prompts with reference images. In this work, we propose Hierarchical Concept-to-Appearance Guidance (CAG), a framework that provides explicit, structured supervision from high-level concepts to fine-grained appearances. At the conceptual level, we introduce a VAE dropout training strategy that randomly omits reference VAE features, encouraging the model to rely more on robust semantic signals from a Visual Language Model (VLM) and thereby promoting consistent concept-level generation in the absence of complete appearance cues. At the appearance level, we integrate the VLM-derived correspondences into a correspondence-aware masked attention module within the Diffusion Transformer (DiT). This module restricts each text token to attend only to its matched reference regions, ensuring precise attribute binding and reliable multi-subject composition. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the multi-subject image generation, substantially improving prompt following and subject consistency.
Yijia Xu, Zihao Wang, Haokun Gui +1
Peking University · Hong Kong University of Science and Technology