Text-to-image diffusion models have achieved remarkable progress in image synthesis, yet can exhibit memorization by closely reproducing individual training examples. Effective mitigation must preserve useful prompt information to guide alternative depictions. We introduce a training-free method that redistributes cross-attention with Gaussian smoothing before reinforcing content-token contributions and attenuating padding contributions, without additional denoiser evaluations. With this intervention, stronger content conditioning can improve prompt alignment at comparable training-image similarity. A local analysis identifies when reinforcement preserves shared value information while redistribution reduces localized attention mass. On Stable Diffusion v1.4 and v2.0, all evaluated smoothing widths lie on the empirical Pareto frontiers for training-image similarity versus both prompt alignment and image preference. A configuration selected on Stable Diffusion reduces template reproduction in DeepFloyd IF without further tuning. These findings support jointly controlling conditioning allocation and strength to generate prompt-consistent alternatives.
Figures & tables
Figure 1: Reinforce a redistributed conditioning pattern. (a) An illustrative attention row shows Gaussian redistribution followed by content reinforcement and padding attenuation, on a shared scale. (b) Recorded outputs illustrate an alternative depiction that retains the prompt context.
Figure 2: Mitigation–quality trade-offs on (a) SD v1.4 and (b) SD v2.0. All recorded operating points are shown; lines connect settings within each method. Lower-right is preferable.
Figure 3: Prompt-consistent alternatives. Comparison with GPA and CS for three prompts, with training references shown for context. Our method retains the green kitchen, fish-patterned blanket, and performance context in alternative compositions.
Figure 4: Reinforcement supports prompt conditioning. Content gain varies with redistribution and suppression retained: (a) prompt alignment (CLIP); (b) training-image similarity (SSCD). Filled markers indicate λw=1 . (c) The guitar itself is blue at higher gain, matching the requested attribute.
Redistribution
Row norm.
SSCD ↓
CLIP ↑
ImageReward ↑
None
No
0.2052
0.2604
−0.3663
Gaussian (ours) , σ=4
No
0.2088
0.2687
−0.2580
Uniform (content)
No
0.2279
0.2751
−0.2086
None
Yes
0.3831
0.2913
0.0467
Gaussian (ours) , σ=4
Yes
0.3312
0.2793
−0.0231
Uniform (content)
Yes
0.3556
0.2875
0.0449
Table 1: Redistribution changes the trade-off under fixed token gains. All intervention rows use λw=2 , λp=0.1 , and full intervention strength ( α=1 ). “None” removes redistribution only. “Gaussian” uses our redistribution operator; “Uniform” acts only within content. Our main method uses no row normalization.
Direction
Norm
SSCD ↓
CLIP ↑
ImageReward ↑
Δ SSCD vs. base
Base
Base
0.5465
0.3028
0.0594
–
Base
Intervention
0.4988
0.2964
0.0150
−0.0477
Intervention
Base
0.3433
0.2826
−0.0179
−0.2032
Intervention
Intervention
0.3458
0.2851
0.0245
−0.2007
Table 2: Suppression persists after restoring guidance magnitude. Direction and norm are controlled at the same latent state.
Figure 5: Preserving EOS balances suppression with context. (a) Our method with and without EOS attenuation. (b) In a separate token-removal example, removing EOS produces a group of people instead of the requested baby, whereas removing padding retains the baby depiction.
IF setting
Template matches / 144
Template SSCD ↓
CLIP ↑
ImageReward ↑
Base
56 (38.89%)
0.1985
0.3045
0.3988
Ours
29 (20.14%)
0.1528
0.2842
0.1058
Table 3: Suppression transfers to DeepFloyd IF. Our SD-selected configuration is transferred without IF-specific tuning.
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Paper notation
Implementation identifier
Ours ( σ=4 , α=1 )
published_s4 / full-strength Gaussian
Positional
abl_position_norm1
Ours ( σ=4 , α=0.625 )
foca_s4_t0p625
Content-uniform
content_uniform_t1p5
Base direction / positional norm
control_base_direction
Positional direction / base norm
control_foca_direction
Appendix
Table 4: Mapping between paper labels and implementation settings.
Figure 6: Reinforcing redistributed attention. Gaussian smoothing precedes content reinforcement and padding attenuation. The original and modified attention are mixed with strength α before value aggregation.
Setting
Redistribution
ρ
α
λw
λp
Norm.
Ours (benchmark)
Global Gaussian, σ∈{1,4,10,100}
–
1
2
0.1
No
Positional
Content positional, σc=1
1
1
2
0.1
Yes
Ours (moderate)
Global Gaussian, σ=4
–
0.625
2
0.1
No
Content-uniform
Content uniform
0.5
1
21.5
0.11.5
Yes
Appendix
Table 5: Settings for the main method and component variations. “Global” excludes SD’s start token. In all controlled SD settings, redistribution acts in the first 5 of 50 steps and weighting acts in all 50. Positional is used for mechanism analysis; our moderate-strength setting and Content-uniform form the fixed-setting transfer comparison. The uniform rows in Table 1 instead use its common gains λw=2 and λp=0.1 .
Redistribution
Row norm.
SSCD ↓
CLIP ↑
ImageReward ↑
None
No
0.2052
0.2604
−0.3663
Gaussian (ours) , σ=1
No
0.2200
0.2737
−0.2041
Gaussian (ours) , σ=4
No
0.2088
0.2687
−0.2580
Content Gaussian ( σc=1 , ρ=1 )
No
0.2224
0.2746
−0.2112
Uniform (content-only, ρ=0.5 )
No
0.2279
0.2751
−0.2086
None
Yes
0.3831
0.2913
0.0467
Appendix
Table 6: Fixed-gain redistribution and normalization results, including the narrower Gaussian kernel and content-only controls. All intervention rows use λw=2 , λp=0.1 , and α=1 . Content-only Gaussian redistributes according to token-position distance and preserves content mass before weighting.
Figure 7: Padding suppression trades alignment for lower similarity. Decreasing λp from 1 to 0 lowers both SSCD and CLIP in the SD v1.4 sweep. Markers follow the recorded settings; the line connects successive settings.
EOS preserved (Ours)
EOS suppressed
σ
SSCD ↓
ImageReward ↑
CLIP ↑
SSCD ↓
ImageReward ↑
CLIP ↑
1
0.2332
−0.350
0.272
0.2112
−0.509
0.265
4
0.2082
−0.413
0.262
0.1899
−0.570
0.257
10
0.1656
−0.637
0.249
0.1518
−0.762
0.243
100
0.1453
−0.721
0.243
0.1382
−0.823
0.238
Appendix
Table 7: Effect of EOS suppression within our method on 500 SD v1.4 prompts, with four images per prompt. Only the first EOS multiplier changes: 1 when preserved and 0.1 when suppressed.
Figure 8: EOS-boundary trade-offs for CLIP and ImageReward. Lines connect the four evaluated widths per variant. Lower-right is preferable; SSCD uses the shared vertical scale.
Figure 9: EOS removal changes prompt content. A token-removal example for Squirrel eating a burger . Removing padding retains the depicted interaction more closely than removing EOS. This is a separate qualitative intervention from the gain-scaling experiment in Figure 8 .
Global smoothing alternative
Strength
SSCD ↓
CLIP ↑
Temperature
τ=1.5
0.2927
0.2722
Temperature
τ=2.0
0.1880
0.2237
Temperature
τ=2.5
0.1143
0.1799
Global uniform
β=0.01
0.1847
0.2583
Global uniform
β=0.05
0.1965
0.2628
Global uniform
β=0.10
0.2212
0.2694
Appendix
Table 8: Global smoothing alternatives on SD v1.4 under the ten-image benchmark protocol. Global uniform smoothing differs from the content-only uniform variant in Table 1 .
Metric
Ours
Content-uniform
Difference [95% CI]
GT SSCD ↓
0.36327
0.34738
−0.01589 [ −0.02435,−0.00787 ]
CLIP ↑
0.28919
0.28701
−0.00219 [ −0.00381,−0.00058 ]
ImageReward ↑
-0.07670
-0.02627
+0.05043 [ +0.02581,+0.07574 ]
High-similarity batch (%) ↓
50.20
45.39
−4.80 [ −7.06,−2.75 ]
Appendix
Table 9: Moderate-strength comparison on 340 prompts and three seeds. Differences are Content-uniform minus Ours, with 95% prompt-bootstrap CIs. High-similarity rates use 1,020 four-image batches per arm; the rate difference is in percentage points.
Type
Metric
Difference
95% CI
TM
GT SSCD
−0.00961
[ −0.01628,−0.00335 ]
TM
CLIP
−0.00330
[ −0.00494,−0.00173 ]
TM
ImageReward
+0.03629
[ +0.01019,+0.06224 ]
TM
High-similarity batch (pp)
−3.28
[ −5.43,−1.39 ]
VM
GT SSCD
−0.03773
[ −0.06754,−0.00951 ]
VM
CLIP
+0.00170
[ −0.00276,+0.00643 ]
Appendix
Table 10: Paired differences (Content-uniform minus Ours) by recorded memorization type: TM, 264 prompts; VM, 76 prompts. Seeds are averaged within each prompt before bootstrapping. Event differences are percentage points.
Seed
Δ SSCD [95% CI]
Δ CLIP
Δ ImageReward
Δ event (pp)
1009
−0.02906 [ −0.03914,−0.01949 ]
−0.00276
+0.04530
−5.88
1013
−0.00290 [ −0.01406,+0.00780 ]
−0.00140
+0.03608
−4.12
1019
−0.01573 [ −0.02475,−0.00718 ]
−0.00240
+0.06992
−4.41
Appendix
Table 11: Per-seed paired differences, Content-uniform minus Ours, over the same 340 prompts. SSCD differences include 95% paired prompt-bootstrap CIs. Δ event is the high-similarity batch-rate difference in percentage points.
IF setting
Template matches / 144
Template SSCD ↓
CLIP ↑
ImageReward ↑
Base
56 (38.89%)
0.1985
0.3045
0.3988
Content-uniform
54 (37.50%)
0.1939
0.2991
0.4556
Ours
29 (20.14%)
0.1528
0.2842
0.1058
Appendix
Table 12: Complete comparison of SD-selected configurations on DeepFloyd IF. Both interventions retain their SD settings. All arms use the same prompts, seeds, and template metrics as Table 3 .
Threshold
Normal flagged (%)
Memorized flagged (%)
Pipeline SSCD
Normal ImageReward
1.429
20.0
80.8
0.2783
0.2345
1.814
5.0
57.8
0.3475
0.2302
2.104
1.0
51.8
0.3622
0.2340
3.736
0.0
32.2
0.4078
0.2366
Appendix
Table 13: Detector threshold sweep with our full-strength configuration. Flag rates are empirical measurements on the stated calibration and evaluation sets.
Method
Recorded setting
SSCD
CLIP
ImageReward
Base
–
0.5439
0.3012
-0.0140
Ours
σ=1
0.2331
0.2773
-0.1868
Ours
σ=4
0.2265
0.2731
-0.2374
Ours
σ=10
0.1838
0.2585
-0.3779
Ours
σ=100
0.1553
0.2484
-0.5188
RTA
tokens=1
0.4794
0.2903
-0.1013
Appendix
Table 14: Complete SD v1.4 aggregate results used in Figure 2 . AIN parameters follow the implementation notation. Values are rounded to four decimal places.
Method
Recorded setting
SSCD
CLIP
ImageReward
Base
–
0.2949
0.2640
-0.4280
Ours
σ=1
0.2090
0.2579
-0.5482
Ours
σ=4
0.1654
0.2208
-1.0782
Ours
σ=10
0.1638
0.2201
-1.1037
Ours
σ=100
0.1617
0.2199
-1.1275
RTA
tokens=1
0.2882
0.2620
-0.4536
Appendix
Table 15: Complete SD v2.0 aggregate results used in Figure 2 . AIN parameters follow the implementation notation. Values are rounded to four decimal places.
Figure 10: Response to semantic prompt edits. Adding Robot or Old changes the subject in our outputs, whereas the shown SD outputs remain close to a common book-cover template. Each row shows one SD output and three outputs from our method for the same edited prompt.
Figure 11: Content reinforcement supports the requested attribute. The complete six-output progression for Blue Guitar – Throw Pillow , ordered by increasing content gain. At higher gain, the guitar printed on the pillow is blue.
While diffusion models excel at generating high-quality images, their tendency to memorize training data poses significant privacy and copyright risks. In this work, we for the first time identify that memorization induces internal numerical instability, often manifesting as visually ``broken'' artifacts. Inspired by stability analysis in numerical methods, we introduce empirical stability regions based on latent update norms to quantitatively characterize stable behavior during generation. Leveraging this, we propose a principled, on-the-fly framework for step-wise detection and adaptive mitigation. Our approach suppresses memorization without altering prompts or guidance, thereby preserving semantic fidelity and image quality. Extensive experiments on Stable Diffusion 1.4 demonstrate that our method achieves an AUC >0.999 detection performance and a 0.0% memorization rate after mitigation with negligible overhead (≈0.01s per image).
Yuanmin Huang, Mi Zhang, Chen Chen +4
Fudan University Shanghai, China · East China University of Science and Technology Shanghai, China
Text-to-image diffusion models are capable of generating high-quality images, but suboptimal pre-trained text representations often result in these images failing to align closely with the given text prompts. Classifier-free guidance (CFG) is a popular and effective technique for improving text-image alignment in the generative process. However, CFG introduces significant computational overhead. In this paper, we present DIstilling CFG by sharpening text Embeddings (DICE) that replaces CFG in the sampling process with half the computational complexity while maintaining similar generation quality. DICE distills a CFG-based text-to-image diffusion model into a CFG-free version by refining text embeddings to replicate CFG-based directions. In this way, we avoid the computational drawbacks of CFG, enabling high-quality, well-aligned image generation at a fast sampling speed. Furthermore, examining the enhancement pattern, we identify the underlying mechanism of DICE that sharpens specific components of text embeddings to preserve semantic information while enhancing fine-grained details. Extensive experiments on multiple Stable Diffusion v1.5 variants, SDXL, and PixArt-α demonstrate the effectiveness of our method. Code is available at https://github.com/zju-pi/dice.
Zhenyu Zhou, Defang Chen, Can Wang +2
Zhejiang University, State Key Laboratory of Blockchain and Data Security · University at Buffalo, State University of New York
Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without retraining. Existing methods either require computationally expensive fine-tuning or rely on style transfer techniques that risk semantic misalignment with textual prompts. We introduce Visual Concept Fusion (VCF), the first method offering dual conditioning on both an image and text prompt at inference time without any concept-specific training. VCF enables visual concept injection into Stable Diffusion by aligning CLIP image features with the text embedding space. VCF consists of three components: (1) a lightweight aligner that maps image tokens to the text embedding manifold using InfoNCE and cross-attention reconstruction losses, (2) a fusion strategy that preserves both textual and visual semantics, and (3) an optional Prompt-Noise Optimization (PNO) module for test-time refinement. Our experiments demonstrate that VCF successfully transfers visual attributes including style, composition, and color palette from reference images while maintaining prompt adherence. Quantitative results show a trade-off between text alignment (CLIP score) and visual correspondence (LPIPS), with VCF outperforming baselines in reference fidelity.