Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color and excess high-frequency detail. To isolate the cause, we fine-tune two models from the same base model with the same DreamBooth recipe, one on real photos of a subject and one on synthetic images of that subject generated by the first. We trace the degradation to classifier-free guidance (CFG). For the model personalized on synthetic images, the angle between the conditional and unconditional noise predictions, and with it the norm of their difference, is much larger than for the model personalized on real photos. This inflation grows toward high frequencies and also appears at other prompts semantically close to the subject, such as its class noun, but not at unrelated ones. We propose ReGain, a training-free correction applied at sampling time that measures how much each frequency band of the guidance is inflated relative to the base model and scales that band down accordingly. ReGain needs no real photos. On Stable Diffusion v1.5, ReGain closes 51-64% of the subject-fidelity gap to the model personalized on real photos, as measured by DINO, DINOv2 and CLIP-I. It also improves subject fidelity on SDXL and SD 3.5 and preserves text alignment on all three backbones.
Figures & tables
Figure 1: Personalizing a model on synthetic images of a subject degrades its fidelity. Left: Mreal is DreamBooth fine-tuned from the base model Mbase on real photographs of the subject, and Msyn is fine-tuned from the same base with an identical recipe on images generated by Mreal . Both are sampled with the same prompt and the same 10 seeds. Right: compared to Mreal , Msyn shows an inflated per-step ∥Δ∥ across denoising steps and an inflated latent power spectrum in its generated images.
Figure 2: (a) Msyn/Mreal ratio of θ , δ and r . (b) Angle excess θMsyn(t)−θMreal(t) for the subject prompt and increasingly semantically distant prompts. 10 seeds mean.
Figure 3: (a) Msyn/Mreal band-energy ratio of Δ per frequency band k and step t ; red marks inflation. (b) θ inflation of the low, mid and high bands over steps. 10 seeds mean.
Figure 4: Overview. (A) Msyn is DreamBooth fine-tuned on synthetic subject images generated by Mreal . (B) The angle θ between the unconditional and conditional predictions is wider for Msyn , so its guidance term Δ=ϵc−ϵ∅ is inflated. (C) ReGain. (i) Once per model, Mbase and Msyn are evaluated at forward-noised training images zt ; the square root of the ratio of their masked band energies e gives the gain g^(k,t) , compressed into the schedule gb(t) (darker blue attenuates more; Figure 5 ). (ii) At each step, Δ is rescaled per band group by gb(t) inside the subject mask m (green) and left unchanged outside it (grey), then scaled by w and added to ϵ∅ , equation 5 .
Figure 5: Gain schedule gb(t) for dog6 from the base model (equation 4 ), one row per band group b .
Figure 6: Qualitative comparison on SD1.5 with DreamBooth. Each row is one subject, prompt and seed shared by the three models. More subjects and settings are in Appendix E .
SD1.5, DreamBooth
SD1.5, DreamBooth-LoRA
Method
CLIP-I
DINO
DINOv2
CLIP-T
CLIP-I
DINO
DINOv2
CLIP-T
Mreal + CFG
0.814
0.681
0.661
0.300
0.791
0.630
0.606
0.306
Msyn + CFG
0.788
0.611
0.611
0.289
0.750
0.540
0.510
0.300
Msyn + S-CFG
0.781
0.605
0.604
0.292
0.744
0.536
0.503
0.302
Msyn + CFG++
0.788
0.606
0.608
0.289
0.748
0.535
0.508
0.299
Msyn + FDG
0.792
0.608
0.614
0.284
0.755
0.538
0.517
0.297
Table 1: Subject and prompt fidelity on DreamBooth, averaged over 30 subjects, 25 prompts and 3 seeds (higher is better). The first row is the reference model Mreal , and every other row samples from Msyn . Baselines, on SD1.5 only: S-CFG ( Shen et al., 2024 ) , CFG++ ( Chung et al., 2025 ) and FDG ( Sadat et al., 2025b ) . Bold marks the best Msyn result. Gap recovered is the part of the drop from Mreal to Msyn , both with CFG, that ReGain wins back. For CLIP-T we report the change from Msyn + CFG instead.
Method
Saturation
Contrast
High band
Mreal + CFG
0.301
0.211
0.061
Msyn + CFG
0.355
0.258
0.069
Msyn + CFG + ReGain
0.309
0.234
0.066
↪ Gap recovered
85%
51%
38%
Table 2: Over-guidance artifacts inside the subject region (SD1.5, DreamBooth). Bold marks the Msyn row closer to Mreal .
Figure 7: Qualitative comparison across fine-tuning forms and base models. Each row is one base model and each half one subject, prompt and seed shared by the three models; nothing is retuned per base model. The SD3.5 row shows the chain of Table 1 .
Table 10
Figure 8: Example generations from the baselines of Table 3 , all at the same seed.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Guidance weight w
Low k
Mid k
High k
All
7.5
4.16
4.71
6.55
5.57
5.0
3.86
4.28
5.80
4.98
3.0
3.50
4.09
5.79
4.88
Appendix
Table 5: Δ inflation of dog6 against the base model, by band group, for Msyn trained on images sampled from Mreal at three guidance weights (root mean square over the group’s bands and all steps at the calibration states, as the Before column of Table 4 ). The 7.5 row is the Msyn used throughout the paper.
Figure 9: Per-step Msyn/Mreal ratio of θ , δ and r for the 30 DreamBooth subjects, as in Figure 2 a. Mean over 5 seeds, shading is the standard error.
Figure 10: Subject mask m during sampling. The mask of equation 6 (red) at steps 0 , 10 , 20 and 40 of a 50 -step plain-CFG trajectory of Msyn , drawn on the clean image predicted at that step, and the mask of the last step on the final sample. Stable Diffusion v1.5, DreamBooth, seed 100 . Each row gives the subject and the scene phrase of its evaluation prompt “a [V] ⟨ class ⟩ …” (Appendix C.1 ).
SD v1.5
SD v1.5
SDXL 1.0
SD 3.5 Medium
DreamBooth
DreamBooth-LoRA
DreamBooth-LoRA
DreamBooth-LoRA
Fine-tuning
Trained weights
full U-Net
U-Net attention
U-Net attention
MMDiT
LoRA rank
–
16
4
4
Learning rate (constant)
5×10−6
10−4
5×10−4
4×10−4
Steps
500
500
500
600
Appendix
Table 6: Fine-tuning and sampling settings. One column per setting of Table 1 . LoRA adapters use a scaling factor α equal to the rank and are merged into the base model weights before sampling. All text encoders are frozen and no prior-preservation loss is used. The synthetic training set is sampled with each base model’s default scheduler and evaluation uses DDIM, except on Stable Diffusion 3.5, which uses its flow-matching Euler scheduler (shift 3.0 ) in both cases.
Figure 11: Estimated gain schedules of three further subjects. The compressed schedule gb(t) at τ=0.05 of the pink sunglasses, cat2 and backpack subjects of Figure 6 , Stable Diffusion v1.5 DreamBooth, in the layout of Figure 5 : one row per band group, one column per sampling step, and color and label giving the gain.
Figure 12: DreamBooth-LoRA, Stable Diffusion v1.5. The block of Figure 6 in the DreamBooth-LoRA setting: the subject’s reference photos, the prompt, and then Mreal , Msyn under plain CFG, and the same Msyn under its estimated gain schedule (ReGain, ours), at one seed and w=7.5 .
Figure 13: DreamBooth-LoRA, SDXL base 1.0. The same block on SDXL at 1024×1024 . Columns and protocol follow Figure 12 .
Figure 14: Additional subjects and prompts. Four further subject and prompt combinations, shown in the same block at the same seed and guidance weight ( w=7.5 , 50 steps): the subject’s reference photos, the prompt, then Mreal , Msyn under plain CFG, and the same Msyn under its estimated gain schedule (ReGain, ours).
Figure 15: Subject-agnostic alternatives, further rows. Four more rows in the layout of Figure 8 , showing Mreal , Msyn under plain CFG at w=7.5 , 5.0 and 3.0 , under APG at w=7.5 , and under ReGain, at the same seed across columns.
School of Artificial Intelligence, Jilin University · 2Shanghai Innovation Institute · 4Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China +1