Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose gated directions enter frozen text-encoder and MMDiT context states as relative residuals. Separating this memory from backbone adaptation enables prompt-selective and multi-concept access without merging model updates. Experiments show that V-Engram broadly matches DreamBooth-LoRA in overall subject fidelity while showing advantages in settings such as contextual subject preservation. Prompt-matched loading retrieves only matched entries, reducing most additional adaptation-state loading for a single-concept query. Qualitative results further demonstrate paired-trigger composition and same-class separation, while prompts without registered entries retain the frozen model's base behavior. Together, these results establish trigger-indexed memory as a modular interface for adding targeted visual evidence without rewriting the generator.
Figures & tables
Figure 2: Overview of V-Engram. Few-shot references and an explicit trigger construct tokenizer-specific, layer-wise Engram entries. Exact token-id matches retrieve entries whose residuals are injected into selected frozen Stable Diffusion 3.5 layers.
Avg. best-ref
Avg. all-pairs
Method
DINOv2 ↑
CLIP-I ↑
DINOv2 ↑
CLIP-I ↑
Zero-shot SD3.5
0.4274
0.6765
0.3501
0.6361
TI-style
0.4919
0.7188
0.4071
0.6813
DB-LoRA ( r=4 )
0.8084
0.8682
0.6911
0.8290
DB-LoRA ( r=8 )
0.8167
0.8875
0.6930
0.8439
V-Engram
0.7896
0.8902
0.6809
0.8463
Table 1: Main comparison over 15 subjects. Both reference aggregations are averaged over five generations and then macro-averaged over subjects.
Figure 3: Qualitative main comparison under the common prompt template “a photo of <trigger> ”. Rows show the personalized clock, fancy boot, and gray sloth plushie; columns show one reference image, zero-shot SD3.5, TI-style / embedding-only personalization, rank-4 DreamBooth-LoRA, and V-Engram.
Mounting setting
DINOv2 ↑
CLIP-I ↑
Encoder only
0.6132
0.7991
N=3
0.7612
0.8499
N=4
0.7814
0.8719
N=5
0.7595
0.8820
N=6
0.7780
0.8894
N=7
0.8068
0.9007
Table 2: Mounting-density ablation. N is the number of MMDiT blocks in addition to encoder sites.
Figure 4: Qualitative mounting-density ablation. N denotes the number of selected MMDiT blocks. Compared with encoder-only injection, mounted variants more consistently preserve the subject’s distinctive shape and appearance, with visible variation across densities.
CLIP-T
CLIP-I
Method
Score ↑
Best-ref ↑
All-pairs ↑
DB-LoRA ( r=4 )
0.2558
0.8264
0.7888
DB-LoRA ( r=8 )
0.2658
0.8105
0.7706
V-Engram
0.2334
0.8717
0.8261
Table 3: Contextual composition under five category-agnostic templates shared by all 15 subjects. Scores are macro-averaged over subjects.
Figure 5: Additional contextual prompts for a personalized cat, shown separately from the five-template protocol in Table 3 . The reference is at left; the upper and lower rows show rank-4 DreamBooth-LoRA and V-Engram, respectively. Exact prompts and additional examples are provided in the supplementary material.
Figure 6: Paired-trigger examples: teapot+vase, cat+toy, and sunglasses+plushie. Columns show two references, V-Engram, and rank-4 DreamBooth-LoRA.
Figure 7: Same-class disambiguation under indexed and name-based triggers for the red-car prompt. Four of the seven dog identities are shown; rows compare V-Engram and rank-4 DreamBooth-LoRA under the two naming schemes.
InsightFace
Method
Best-ref ↑
All-pairs ↑
Valid
Zero-shot Stable Diffusion 3.5
0.2344
0.1824
75/75
DB-LoRA ( r=4 )
0.4618
0.4014
75/75
V-Engram
0.6084
0.5415
75/75
Table 4: InsightFace identity similarity on 15 familiar identities. Scores are averaged over generations and macro-averaged over identities.
Figure 8: Qualitative familiar-identity refinement for three of the 15 evaluated identities. Columns show a held-out reference, zero-shot Stable Diffusion 3.5, rank-4 DreamBooth-LoRA, and V-Engram.
Current personalization methods for generative vision models typically encode new concepts through continuous adapters or weight updates, yet provide limited control over whether and when a concept should be retrieved. In this work, we introduce Tiny-Engram, a compact trigger-indexed concept table that gives visual memories an explicit lexical address and activation boundary inside frozen image and video generators. Tiny-Engram parameterizes each concept as a small set of memory entries indexed by registered n-gram matches, which modulate text-encoder hidden states only within the matched trigger region. Outside this lexical support, the conditioning pathway is identical to that of the frozen base model. Across both single-encoder latent diffusion and multi-encoder diffusion-transformer backbones, this formulation binds a rare trigger phrase to a target identity while preserving compositional control from the surrounding prompt. We further evaluate the same table-based memory in a text-conditioned video generation setting, where the trigger path reliably alters the generated subject but fine-grained identity persistence across held-out video prompts remains limited. Taken together, these results suggest that small, explicitly addressed concept tables are a practical route to modular visual personalization, with strongest evidence in image generation. For video diffusion, the remaining gap points to a broader requirement: temporally stable identity likely depends on tighter coupling between text-side memory and the evolving visual state, motivating future work on memory injection beyond the text-conditioning interface.
Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without retraining. Existing methods either require computationally expensive fine-tuning or rely on style transfer techniques that risk semantic misalignment with textual prompts. We introduce Visual Concept Fusion (VCF), the first method offering dual conditioning on both an image and text prompt at inference time without any concept-specific training. VCF enables visual concept injection into Stable Diffusion by aligning CLIP image features with the text embedding space. VCF consists of three components: (1) a lightweight aligner that maps image tokens to the text embedding manifold using InfoNCE and cross-attention reconstruction losses, (2) a fusion strategy that preserves both textual and visual semantics, and (3) an optional Prompt-Noise Optimization (PNO) module for test-time refinement. Our experiments demonstrate that VCF successfully transfers visual attributes including style, composition, and color palette from reference images while maintaining prompt adherence. Quantitative results show a trade-off between text alignment (CLIP score) and visual correspondence (LPIPS), with VCF outperforming baselines in reference fidelity.
Text-to-image diffusion models have achieved remarkable progress in image synthesis, yet can exhibit memorization by closely reproducing individual training examples. Effective mitigation must preserve useful prompt information to guide alternative depictions. We introduce a training-free method that redistributes cross-attention with Gaussian smoothing before reinforcing content-token contributions and attenuating padding contributions, without additional denoiser evaluations. With this intervention, stronger content conditioning can improve prompt alignment at comparable training-image similarity. A local analysis identifies when reinforcement preserves shared value information while redistribution reduces localized attention mass. On Stable Diffusion v1.4 and v2.0, all evaluated smoothing widths lie on the empirical Pareto frontiers for training-image similarity versus both prompt alignment and image preference. A configuration selected on Stable Diffusion reduces template reproduction in DeepFloyd IF without further tuning. These findings support jointly controlling conditioning allocation and strength to generate prompt-consistent alternatives.