Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose gated directions enter frozen text-encoder and MMDiT context states as relative residuals. Separating this memory from backbone adaptation enables prompt-selective and multi-concept access without merging model updates. Experiments show that V-Engram broadly matches DreamBooth-LoRA in overall subject fidelity while showing advantages in settings such as contextual subject preservation. Prompt-matched loading retrieves only matched entries, reducing most additional adaptation-state loading for a single-concept query. Qualitative results further demonstrate paired-trigger composition and same-class separation, while prompts without registered entries retain the frozen model's base behavior. Together, these results establish trigger-indexed memory as a modular interface for adding targeted visual evidence without rewriting the generator.
Figures & tables
Figure 2: Overview of V-Engram. Few-shot references and an explicit trigger construct tokenizer-specific, layer-wise Engram entries. Exact token-id matches retrieve entries whose residuals are injected into selected frozen Stable Diffusion 3.5 layers.
Avg. best-ref
Avg. all-pairs
Method
DINOv2 ↑
CLIP-I ↑
DINOv2 ↑
CLIP-I ↑
Zero-shot SD3.5
0.4274
0.6765
0.3501
0.6361
TI-style
0.4919
0.7188
0.4071
0.6813
DB-LoRA ( r=4 )
0.8084
0.8682
0.6911
0.8290
DB-LoRA ( r=8 )
0.8167
0.8875
0.6930
0.8439
V-Engram
0.7896
0.8902
0.6809
0.8463
Table 1: Main comparison over 15 subjects. Both reference aggregations are averaged over five generations and then macro-averaged over subjects.
Figure 3: Qualitative main comparison under the common prompt template “a photo of <trigger> ”. Rows show the personalized clock, fancy boot, and gray sloth plushie; columns show one reference image, zero-shot SD3.5, TI-style / embedding-only personalization, rank-4 DreamBooth-LoRA, and V-Engram.
Mounting setting
DINOv2 ↑
CLIP-I ↑
Encoder only
0.6132
0.7991
N=3
0.7612
0.8499
N=4
0.7814
0.8719
N=5
0.7595
0.8820
N=6
0.7780
0.8894
N=7
0.8068
0.9007
Table 2: Mounting-density ablation. N is the number of MMDiT blocks in addition to encoder sites.
Figure 4: Qualitative mounting-density ablation. N denotes the number of selected MMDiT blocks. Compared with encoder-only injection, mounted variants more consistently preserve the subject’s distinctive shape and appearance, with visible variation across densities.
CLIP-T
CLIP-I
Method
Score ↑
Best-ref ↑
All-pairs ↑
DB-LoRA ( r=4 )
0.2558
0.8264
0.7888
DB-LoRA ( r=8 )
0.2658
0.8105
0.7706
V-Engram
0.2334
0.8717
0.8261
Table 3: Contextual composition under five category-agnostic templates shared by all 15 subjects. Scores are macro-averaged over subjects.
Figure 5: Additional contextual prompts for a personalized cat, shown separately from the five-template protocol in Table 3 . The reference is at left; the upper and lower rows show rank-4 DreamBooth-LoRA and V-Engram, respectively. Exact prompts and additional examples are provided in the supplementary material.
Figure 6: Paired-trigger examples: teapot+vase, cat+toy, and sunglasses+plushie. Columns show two references, V-Engram, and rank-4 DreamBooth-LoRA.
Figure 7: Same-class disambiguation under indexed and name-based triggers for the red-car prompt. Four of the seven dog identities are shown; rows compare V-Engram and rank-4 DreamBooth-LoRA under the two naming schemes.
InsightFace
Method
Best-ref ↑
All-pairs ↑
Valid
Zero-shot Stable Diffusion 3.5
0.2344
0.1824
75/75
DB-LoRA ( r=4 )
0.4618
0.4014
75/75
V-Engram
0.6084
0.5415
75/75
Table 4: InsightFace identity similarity on 15 familiar identities. Scores are averaged over generations and macro-averaged over identities.
Figure 8: Qualitative familiar-identity refinement for three of the 15 evaluated identities. Columns show a held-out reference, zero-shot Stable Diffusion 3.5, rank-4 DreamBooth-LoRA, and V-Engram.