Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color and excess high-frequency detail. To isolate the cause, we fine-tune two models from the same base model with the same DreamBooth recipe, one on real photos of a subject and one on synthetic images of that subject generated by the first. We trace the degradation to classifier-free guidance (CFG). For the model personalized on synthetic images, the angle between the conditional and unconditional noise predictions, and with it the norm of their difference, is much larger than for the model personalized on real photos. This inflation grows toward high frequencies and also appears at other prompts semantically close to the subject, such as its class noun, but not at unrelated ones. We propose ReGain, a training-free correction applied at sampling time that measures how much each frequency band of the guidance is inflated relative to the base model and scales that band down accordingly. ReGain needs no real photos. On Stable Diffusion v1.5, ReGain closes 51-64% of the subject-fidelity gap to the model personalized on real photos, as measured by DINO, DINOv2 and CLIP-I. It also improves subject fidelity on SDXL and SD 3.5 and preserves text alignment on all three backbones.
Figures & tables
Figure 1: Personalizing a model on synthetic images of a subject degrades its fidelity. Left: Mreal is DreamBooth fine-tuned from the base model Mbase on real photographs of the subject, and Msyn is fine-tuned from the same base with an identical recipe on images generated by Mreal . Both are sampled with the same prompt and the same 10 seeds. Right: compared to Mreal , Msyn shows an inflated per-step ∥Δ∥ across denoising steps and an inflated latent power spectrum in its generated images.
Figure 2: (a) Msyn/Mreal ratio of θ , δ and r . (b) Angle excess θMsyn(t)−θMreal(t) for the subject prompt and increasingly semantically distant prompts. 10 seeds mean.
Figure 3: (a) Msyn/Mreal band-energy ratio of Δ per frequency band k and step t ; red marks inflation. (b) θ inflation of the low, mid and high bands over steps. 10 seeds mean.
Figure 4: Overview. (A) Msyn is DreamBooth fine-tuned on synthetic subject images generated by Mreal . (B) The angle θ between the unconditional and conditional predictions is wider for Msyn , so its guidance term Δ=ϵc−ϵ∅ is inflated. (C) ReGain. (i) Once per model, Mbase and Msyn are evaluated at forward-noised training images zt ; the square root of the ratio of their masked band energies e gives the gain g^(k,t) , compressed into the schedule gb(t) (darker blue attenuates more; Figure 5 ). (ii) At each step, Δ is rescaled per band group by gb(t) inside the subject mask m (green) and left unchanged outside it (grey), then scaled by w and added to ϵ∅ , equation 5 .
Figure 5: Gain schedule gb(t) for dog6 from the base model (equation 4 ), one row per band group b .
Figure 6: Qualitative comparison on SD1.5 with DreamBooth. Each row is one subject, prompt and seed shared by the three models. More subjects and settings are in Appendix E .
SD1.5, DreamBooth
SD1.5, DreamBooth-LoRA
Method
CLIP-I
DINO
DINOv2
CLIP-T
CLIP-I
DINO
DINOv2
CLIP-T
Mreal + CFG
0.814
0.681
0.661
0.300
0.791
0.630
0.606
0.306
Msyn + CFG
0.788
0.611
0.611
0.289
0.750
0.540
0.510
0.300
Msyn + S-CFG
0.781
0.605
0.604
0.292
0.744
0.536
0.503
0.302
Msyn + CFG++
0.788
0.606
0.608
0.289
0.748
0.535
0.508
0.299
Msyn + FDG
0.792
0.608
0.614
0.284
0.755
0.538
0.517
0.297
Table 1: Subject and prompt fidelity on DreamBooth, averaged over 30 subjects, 25 prompts and 3 seeds (higher is better). The first row is the reference model Mreal , and every other row samples from Msyn . Baselines, on SD1.5 only: S-CFG ( Shen et al., 2024 ) , CFG++ ( Chung et al., 2025 ) and FDG ( Sadat et al., 2025b ) . Bold marks the best Msyn result. Gap recovered is the part of the drop from Mreal to Msyn , both with CFG, that ReGain wins back. For CLIP-T we report the change from Msyn + CFG instead.
Method
Saturation
Contrast
High band
Mreal + CFG
0.301
0.211
0.061
Msyn + CFG
0.355
0.258
0.069
Msyn + CFG + ReGain
0.309
0.234
0.066
↪ Gap recovered
85%
51%
38%
Table 2: Over-guidance artifacts inside the subject region (SD1.5, DreamBooth). Bold marks the Msyn row closer to Mreal .
Figure 7: Qualitative comparison across fine-tuning forms and base models. Each row is one base model and each half one subject, prompt and seed shared by the three models; nothing is retuned per base model. The SD3.5 row shows the chain of Table 1 .
Table 10
Figure 8: Example generations from the baselines of Table 3 , all at the same seed.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Guidance weight w
Low k
Mid k
High k
All
7.5
4.16
4.71
6.55
5.57
5.0
3.86
4.28
5.80
4.98
3.0
3.50
4.09
5.79
4.88
Appendix
Table 5: Δ inflation of dog6 against the base model, by band group, for Msyn trained on images sampled from Mreal at three guidance weights (root mean square over the group’s bands and all steps at the calibration states, as the Before column of Table 4 ). The 7.5 row is the Msyn used throughout the paper.
Figure 9: Per-step Msyn/Mreal ratio of θ , δ and r for the 30 DreamBooth subjects, as in Figure 2 a. Mean over 5 seeds, shading is the standard error.
Figure 10: Subject mask m during sampling. The mask of equation 6 (red) at steps 0 , 10 , 20 and 40 of a 50 -step plain-CFG trajectory of Msyn , drawn on the clean image predicted at that step, and the mask of the last step on the final sample. Stable Diffusion v1.5, DreamBooth, seed 100 . Each row gives the subject and the scene phrase of its evaluation prompt “a [V] ⟨ class ⟩ …” (Appendix C.1 ).
SD v1.5
SD v1.5
SDXL 1.0
SD 3.5 Medium
DreamBooth
DreamBooth-LoRA
DreamBooth-LoRA
DreamBooth-LoRA
Fine-tuning
Trained weights
full U-Net
U-Net attention
U-Net attention
MMDiT
LoRA rank
–
16
4
4
Learning rate (constant)
5×10−6
10−4
5×10−4
4×10−4
Steps
500
500
500
600
Appendix
Table 6: Fine-tuning and sampling settings. One column per setting of Table 1 . LoRA adapters use a scaling factor α equal to the rank and are merged into the base model weights before sampling. All text encoders are frozen and no prior-preservation loss is used. The synthetic training set is sampled with each base model’s default scheduler and evaluation uses DDIM, except on Stable Diffusion 3.5, which uses its flow-matching Euler scheduler (shift 3.0 ) in both cases.
Figure 11: Estimated gain schedules of three further subjects. The compressed schedule gb(t) at τ=0.05 of the pink sunglasses, cat2 and backpack subjects of Figure 6 , Stable Diffusion v1.5 DreamBooth, in the layout of Figure 5 : one row per band group, one column per sampling step, and color and label giving the gain.
Figure 12: DreamBooth-LoRA, Stable Diffusion v1.5. The block of Figure 6 in the DreamBooth-LoRA setting: the subject’s reference photos, the prompt, and then Mreal , Msyn under plain CFG, and the same Msyn under its estimated gain schedule (ReGain, ours), at one seed and w=7.5 .
Figure 13: DreamBooth-LoRA, SDXL base 1.0. The same block on SDXL at 1024×1024 . Columns and protocol follow Figure 12 .
Figure 14: Additional subjects and prompts. Four further subject and prompt combinations, shown in the same block at the same seed and guidance weight ( w=7.5 , 50 steps): the subject’s reference photos, the prompt, then Mreal , Msyn under plain CFG, and the same Msyn under its estimated gain schedule (ReGain, ours).
Figure 15: Subject-agnostic alternatives, further rows. Four more rows in the layout of Figure 8 , showing Mreal , Msyn under plain CFG at w=7.5 , 5.0 and 3.0 , under APG at w=7.5 , and under ReGain, at the same seed across columns.
Recent subject-to-image models have achieved impressive progress in personalized image generation, yet they still struggle to preserve fine-grained subject-specific details. A major reason is the lack of high-quality fine-grained identity supervision: real paired data are expensive to collect, while synthesized training pairs often preserve only coarse subject appearance and fail to capture subtle subject-specific details. In this work, we propose CopyCat, a lightweight model-refinement framework that improves fine-grained subject consistency within only a few seconds. CopyCat performs a one-time refinement of a pretrained subject-to-image model by attaching a lightweight Fine-grained Consistency LoRA (FCLoRA) and optimizing it using a single proxy image, which is used as both the conditioning image and the reconstruction target. This exact self-reconstruction objective substantially simplifies the optimization task, enabling effective fine-grained refinement within only a few seconds. The refinement is performed only once; the resulting model can be directly applied to diverse unseen reference subjects and prompts without further subject-specific optimization. We further revisit subject-to-image LoRA training in double-stream diffusion transformers and find that adapting only the visual stream consistently improves subject consistency. Extensive experiments on DreamBench and XVerseBench demonstrate consistent improvements in fine-grained subject consistency across representative subject-to-image models under both single- and multi-subject settings.
Peng Zheng, Ruiqi Liu, Rui Ma +1
School of Artificial Intelligence, Jilin University · 2Shanghai Innovation Institute · 4Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China +1
Sampling latency compounds in diffusion workflows, where users generate and discard many candidates before keeping one. Surprisingly, the poor outputs of standard few-step samplers do not reflect a lack of reconstruction capacity: by optimizing only the initial noise, frozen 3-4-step samplers can closely reproduce their corresponding full-step outputs. Building on this finding, we learn corrections to the initial noise and denoising updates using endpoint supervision, improving correspondence with full-step outputs generated from the same noise and prompt. The resulting previews allow users to screen candidates cheaply and reserve full-step generation for promising ones. Input correction also transfers across sampling budgets without retraining. Experiments show substantial improvements in reference fidelity, including 53-78% lower reconstruction MSE than retrained LD3 on unconditional benchmarks, alongside improved ranking preservation and candidate selection on SD1.5, SDXL, and FLUX.1-dev.
Jing Jia, Sifan Liu, Guanyang Wang
Department of Computer Science, Rutgers University · Department of Statistical Science, Duke University · Department of Statistics, Rutgers University
Text-to-image diffusion models have achieved remarkable progress in image synthesis, yet can exhibit memorization by closely reproducing individual training examples. Effective mitigation must preserve useful prompt information to guide alternative depictions. We introduce a training-free method that redistributes cross-attention with Gaussian smoothing before reinforcing content-token contributions and attenuating padding contributions, without additional denoiser evaluations. With this intervention, stronger content conditioning can improve prompt alignment at comparable training-image similarity. A local analysis identifies when reinforcement preserves shared value information while redistribution reduces localized attention mass. On Stable Diffusion v1.4 and v2.0, all evaluated smoothing widths lie on the empirical Pareto frontiers for training-image similarity versus both prompt alignment and image preference. A configuration selected on Stable Diffusion reduces template reproduction in DeepFloyd IF without further tuning. These findings support jointly controlling conditioning allocation and strength to generate prompt-consistent alternatives.