Personalizing text-to-image diffusion models extends pretrained models to represent novel user-specific concepts from only a few reference images. However, learning a new concept while building on the prior knowledge of the pretrained model remains a key challenge. When personalization focuses on learning the target concept, the model tends to overfit the reference examples and degrade its general capability. In contrast, emphasizing prior preservation can hinder capturing distinctive personalized attributes. In this paper, we address this challenge by viewing a personalized concept as an underrepresented concept whose semantic counterpart is well represented in the pretrained model. Rather than treating the learning of a new concept and prior preservation as separate objectives, we reformulate them as a single anchored learning problem. We therefore introduce Semantic Anchoring Personalization (SAP), which keeps concept learning grounded in the pretrained semantic structure while capturing subject-specific attributes. The proposed objective offers a simple yet effective formulation that can be applied across different model backbones without architectural modifications or auxiliary networks. Extensive experiments across various settings demonstrate that SAP achieves a better balance between subject fidelity and text-image alignment than baseline methods. Further ablation studies validate the contribution of semantic anchoring to personalization.
Figures & tables
Figure 1: Conceptual illustration of our method, Semantic Anchoring Personalization (SAP). Pretrained semantics provide stable guidance, whereas guidance in newly introduced regions remains unstable. Our approach anchors personalization to pretrained semantics, enabling stable guidance as the model expands toward novel concepts.
Figure 2: L2 distance between the personalized subject prediction and its pretrained class-level counterpart throughout the personalization process.
Figure 3: Visual comparison of few-shot personalization and encoder-based methods on SD1.5.
Table 4
Figure 4: Qualitative comparison on SDXL and SD3 backbones among DreamBooth-LoRA, Beyond-Finetuning, and our method.
Figure 6Figure 7
Group
Setting
CLIP-I ↑
CLIP-T ↑
DINO ↑
dog
dog → dog (Ours)
0.8108
0.3059
0.6230
dog → cat
0.8076
0.3031
0.6170
dog → animal
0.8027
0.3039
0.6130
dog → shark
0.7965
0.3026
0.5918
teapot
teapot → teapot (Ours)
0.8329
0.3400
0.5002
teapot → kettle
0.8024
0.3318
0.4488
Table 3: Ablation study on anchor semantics.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Method
CLIP-I ↑
CLIP-T ↑
DINO ↑
SD1.5
Finetuning Anchoring
0.8112
0.2870
0.6563
Pretrained Anchoring (Ours)
0.7833
0.3064
0.6051
SDXL
Finetuning Anchoring
0.8407
0.2838
0.7109
Pretrained Anchoring (Ours)
0.8005
0.3154
0.6661
SD3
Finetuning Anchoring
0.6182
0.2548
0.2685
Pretrained Anchoring (Ours)
0.7579
0.3187
0.5679
Appendix
Table 4: Comparison on SD1.5, SDXL, SD3 using two different anchoring strategies.
Method
CLIP-I ↑
CLIP-T ↑
DINO ↑
Ours
0.7833
0.3064
0.6051
Pretrained Unconditional Guidance
0.7869
0.2791
0.6050
Finetuned Anchored Guidance
0.7250
0.2497
0.4935
Pretrained Anchored Guidance
0.7888
0.2506
0.6049
Appendix
Table 5: Ablation of inference-time guidance strategies on SD1.5 under the model trained with our semantic anchoring objective.
Method
Identity Preserv.
Prompt Align.
Visual Quality
Overall
Ours
64.00
74.50
65.17
78.50
DreamBooth
22.67
12.50
18.00
14.33
DreamBooth-LoRA
1.33
10.50
6.83
4.50
IP-Adapter
12.00
2.50
10.00
2.67
Appendix
Table 6: User study results across four criteria (all values in %).
Figure 9: Additional visual comparison of few-shot personalization and encoder-based methods on the SD1.5 backbone.
Figure 10: Additional visual comparison of few-shot personalization and encoder-based methods on the SD1.5 backbone, including representative failure cases.
Figure 11: Additional qualitative comparison on SD1.5, SDXL, and SD3 backbones among DreamBooth-LoRA, Beyond-Finetuning, and our method.
Figure 12: Comparison of qualitative results under different weighting parameters w .
Figure 13: Comparison of qualitative results using two different anchoring strategies.
Figure 14: Visual comparison of different guidance strategies during inference.
While text-to-image diffusion models achieve impressive visual quality, they frequently struggle to maintain precise alignment with complex compositional prompts. An effective strategy is to improve the inference process of diffusion models, thereby better leveraging their pretrained priors to address misalignment. Existing training-free methods can be divided into two categories. The first category focuses on improving the randomly sampled initial noise, either performing costly search over noise pools or manipulating sampled noise without ensuring reliable semantic injection. The second category focuses on improving the denoising trajectory, lacking explicit mechanisms to timely diagnose and correct semantic errors. we propose \textbf{AnchorSteer}, a training-free framework that exerts fine-grained control over \textbf{both initialization} and \textbf{the denoising trajectory}. AnchorSteer consists of two synergistic components: \textbf{Semantic Anchoring} replaces uninformative Gaussian noise with text-aligned initializations via CLIP-based prior extraction and a novel Latent-Prior Score Distillation Sampling (LP-SDS) objective. Specifically, LP-SDS distills CLIP visual priors into the knowledge distribution of diffusion models, mitigating the domain gap between CLIP-based priors and diffusion-based priors. \textbf{Reflective Steering} transforms passive denoising with an active Think--Erase--Retouch loop that enables mid-generation self-correction. It leverages VLM-based diagnosis to detect semantic deviations and performs targeted latent refinement to suppress erroneous content and recover missing attributes. Extensive experiments on GenEval and T2I-CompBench++ demonstrate that AnchorSteer consistently outperforms existing baselines in text--image alignment while preserving high visual quality.
Personalizing text-to-image diffusion models to render several specific subjects in a coherent image remains challenging: the model must preserve each subject's identity while keeping the scene spatially and visually coherent. Methods that fuse independently trained concept adapters in a shared weight space (via federated averaging, gradient fusion, or orthogonality constraints) suffer from identity confusion and style bleeding. In this work, we show that composing concepts as separate image layers, instead of merging their adapters in a shared weight space, avoids parameter-level interference. We introduce LILAC, a framework that composes independently trained low-rank adapters at inference time: each subject is conditioned on the frozen composite of previously placed subjects, with exactly one adapter active at a time, therefore identities never interfere at the parameter level. LILAC composes the adapters without joint training, scales linearly with the number of concepts, and is backbone-agnostic. Under the Orthogonal Adaptation protocol, LILAC applied on Qwen-Image-Edit+Qwen-Image-Layered reaches an ArcFace detection rate of 0.861. Code is available at https://github.com/marianlupascu/LILAC.
Marian Lupascu, Sebastian Ripa, Mihai Trascau +2
Adobe Research, Romania · International Computer High School of Bucharest, Romania
Text-to-image personalization aims to generate a user-provided subject in novel scenes described by text. However, most existing methods encode subject identity (fidelity) and context (editability) through the same conditioning pathway, forcing the two to compete for attention-map resources. We refer to this phenomenon as conditioning entanglement and show that it induces a fidelity-editability trade-off. We further provide causal evidence by replacing the target subject token with a generic subject token, which produces shifts in attention allocation and corresponding changes in context adherence. To this end, we propose Decoupled Guidance (DeGu), a plug-and-play framework that routes subject identity and scene context through two independent guidance streams. We further introduce a spatial mixing mechanism that dynamically fuses these streams, ensuring each operates within its semantically relevant region without interference. Furthermore, DeGu can be readily applied to existing personalization methods without modifying the underlying backbone models, consistently improving the overall personalization performance while enabling inference-time control over the fidelity-editability balance, across diverse methods and backbones, including flow-matching Diffusion Transformers (DiTs).
Seongmin Kim, Kyucheol Shin, Heesun Jung +2
Department of Artificial Intelligence Hanyang University · Department of Data Science Hanyang University