Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive as the number and resolution of references grow. We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms. Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues. Condition routing and attention routing align reference tokens with their assigned target regions and restrict cross-reference interactions, while allowing selective reference access beyond region boundaries for scene integration. We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references. After many-reference fine-tuning, RefRoute achieves an overall Weighted-Ref-VIEScore of 36.06 on ManyRef100, compared with 8.88 for FLUX.2-Klein-9B. Separate inference-cost evaluations show substantially slower latency growth as the reference count increases: at 16 references, our 50-step and 4-step configurations achieve 18.3× and 14.2× speedups over their corresponding FLUX baselines, respectively. These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.
Figures & tables
Figure 1: Demo of RefRoute. Our method enables efficient many-reference generation while preserving fine appearance details and spatial correspondence. With 16 references, we achieve an 18.3 × speedup over FLUX.2-Klein-9B at 50 steps.
Figure 2: Overview of RefRoute. Each reference is downsampled and VAE-encoded into compact tokens, while a residual encoder extracts complementary features from the full-resolution image. These features are added to the corresponding reference embeddings after the DiT input projection. The augmented embeddings are jointly processed with text and noisy target embeddings through condition-specific adapters and attention routing.
Method
Object
Person
HOI
De&Re
Overall
Gemini-3-Pro-Image-Preview
50.59
54.75
50.21
52.13
51.76
GPT-Image-1.5
56.66
46.16
52.35
48.46
50.60
Gemini-2.5-Flash-Image
48.01
41.79
49.64
49.44
47.83
Qwen-Image-MICo
52.38
21.11
34.95
37.42
35.86
BAGEL-MICo
38.98
28.45
25.30
44.51
34.41
OmniGen2-MICo
46.26
22.85
32.18
36.82
33.82
Table 1: MICo-Bench results. Completed evaluations cover all 897 cases at native 1024×1024 resolution. Columns report category-level and overall Weighted-Ref-VIEScore ( 0 – 100 , higher is better). Rows above the first rule are reported results; FLUX.2 rows and ours are evaluated by us.
Figure 3: Qualitative comparison. Top: MICo-Bench, with few references; columns show the references, the target, RefRoute, open-source baselines (Qwen-Image-MICo, FLUX.2-Klein-9B), and closed-source systems (GPT-Image, Gemini). Bottom: ManyRef100, with 10–17 references and assigned regions overlaid.
Method
Overall
Human
Object
Mixed
W
SC
PQ
FLUX.2-Klein-Base-9B
4.18
10.79
0.09
2.28
0.268
1.470
6.949
FLUX.2-Klein-9B
8.88
17.07
4.56
5.98
0.375
2.679
7.319
GPT-Image-1
27.92
10.31
55.95
20.09
0.574
5.613
7.763
Ours-MICo
17.32
13.60
36.10
6.02
0.639
3.758
6.171
Ours-ManyRef
36.06
28.10
57.37
26.04
0.709
6.428
7.583
Table 2: Comparison on ManyRef100. We evaluate 1024×1024 generation with many references on our benchmark. Ours-ManyRef achieves the best overall and per-category scores; without many-reference fine-tuning, Ours-MICo outperforms both FLUX.2 baselines.
Figure 4: Inference cost versus reference count. Compared with FLUX.2-Klein baselines and an OminiControl2 reference-branch implementation, RefRoute exhibits slower latency growth and nearly constant peak allocated memory across the measured configurations.
Table 3: Residual source and injection space. Left: Stage 1 reconstruction on 100 images. Right: ManyRef100 after a matched Stage 2 budget of 10k steps. Human and mixed cells are ArcFace face cosine / mIoU / RRBA; object cells are DINO / mIoU / RRBA. Bold is best in each panel.
Figure 5: Strict spatial masking vs. dynamic top-1 routing with the same checkpoint, seed, and prompt. Insets highlight reflections and ground contact outside assigned regions.
Count MAE ↓
Setting
Overall
Human
Mix
w/o Stage 1
1.643
1.300
1.900
w/ Stage 1
1.057
1.033
1.075
Table 4: Effect of Stage 1 pretraining. Person-count MAE after a matched Stage 2 budget of 10k steps. Lower is better; best results are in bold.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Schematic of the object and person construction branches. The diagram shows target synthesis, decomposition into boxes, masks, and references (plus person poses), followed by verification. The illustrated images explain the flow; corpus counts and quality limits are specified in Appendix A.1 .
Group
Models and use
Ours
Ours-MICo (40k); Ours-ManyRef (high- K continuation of Ours-MICo ). Evaluated by us on MICo-Bench and ManyRef100 at 10242 .
MICo-tuned open models
reported MICo-Bench scores.
Native open models
FLUX.2-Klein-9B (evaluated by us); OmniGen2 and Qwen-Image-Edit (reported MICo-Bench scores).
OminiControl2 mechanism
Released reference-branch code reproduced on our backbone; no updated public weights. 4- and 50-step cost only.
Appendix
Table 5: Models and baseline configurations.
Figure 7: Stage 1 output. Each reference is encoded as a 16×16 grid plus a pixel residual. After Stage 1, the model already preserves facial identity and garment details. Target is the ground-truth image. Bottom: zoomed crops of the red boxes.
Stage 1 (T2I)
Stage 2 (joint training)
Stage 3 (distillation)
Initialized from
Backbone weights
Stage 1 compressor + context LoRA
Trained teacher
Objective
LFM (text-to-image)
LFM (multi-ref, routed)
Distillation following TBSM [ Sun et al., 2026 ]
Trainable
Compressor, context LoRA
Compressor, context + spatial LoRA, layout stream
LoRA adapters
LoRA rank / LR
128/ 10−4
128/ 10−4
128/ 10−5
Training steps
40k
40k (main); 60k total (high- K )
20k
Batch / GPUs
1×4 / 4
1×8 / 8
1×8 / 8
Appendix
Table 6: Training configuration by stage. Stage 2 uses 40k steps for the main experiments. The high- K variant continues training on RefRoute-Data for 20k additional steps (60k total). Stage 3 follows TBSM [ Sun et al., 2026 ] for distillation.
Figure 8: Additional qualitative results on MICo-Bench.
Figure 9: Additional many-reference results with 11 to 14 references.
Reference-based diffusion models enable highly controllable image generation by leveraging elements from input images to guide prompt-driven synthesis. However, these models are computationally expensive in runtime, and their cost scales severely with the number of input references. While the efficiency of diffusion models has been extensively studied in the context of prompt-driven generation, it remains largely under-explored in the realm of reference-based models. This setting presents unique challenges not addressed by methods focusing solely on generation. In particular, the wasteful representation of references as dense token grids offers significant opportunities for improvement. In this work, we present Sparse Context, a method for constructing sparse reference representations by retaining only a reduced subset of reference tokens. We observe that even without modifying the model, dropping a significant portion of reference tokens at inference time largely preserves its generation capabilities. To fully realize this potential, we fine-tune the model with random token dropping at varying ratios, encouraging robustness to partial reference representations. Crucially, this training strategy decouples the model from any specific token selection rule, allowing flexible control at inference time. At inference time, instead of random dropping, we apply task-aware token selection strategies that prioritize the most informative regions of the reference images, adapting the token budget to the input and task requirements. Extensive experiments show our method achieves a 4x increase in inference speed for multi-reference generation and an 2x for single reference generation. Importantly, this efficiency is achieved without compromising visual quality across both spatially-aligned editing and subject-driven generation.
Rishubh Parihar, Ayush Raina, R. Venkatesh Babu +1
In-context diffusion transformers concatenate instruction, target, and reference tokens into a single sequence for joint attention. Reference-side computation must therefore be repeated at every denoising step, with the cost growing rapidly as more references are added. Decoupling reference tokens from the target enables exact key-value reuse across denoising steps, but prevents the references from attending to the instruction, degrading instruction following and reference fidelity. This trade-off cannot be resolved through attention-mask design alone. We introduce AnchorCache, a parameter-free token-layout and attention-mask co-design that inserts static text anchors. These anchors condition the reference representations on the instruction during cache construction, after which the resulting reference keys and values can be reused exactly across denoising steps. To recover the quality initially lost through this structural conversion, we apply teacher-forced velocity distillation followed by a short on-policy stage that queries the teacher at student-visited states. To our knowledge, this is the first use of on-policy distillation for architectural recovery in diffusion models. Across benchmarks spanning image, speech, and video generation, AnchorCache matches full-attention quality. Its efficiency gains increase with the reference-context size, reaching a 6.40x speedup in diffusion transformer inference.
Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enhanced diffusion models commonly rely on decoupled visual conditioning: semantic ViT features are processed by the VLM for instruction understanding, whereas appearance-rich VAE features are injected later into the diffusion backbone. Despite its intuitive design, this separation makes it difficult for the model to associate each semantically grounded subject with visual details from the correct reference image. As a result, the model may recognize which subject is being referred to, but fail to preserve its identity and fine-grained appearance, leading to attribute leakage and cross-reference confusion in complex multi-reference settings. To address this issue, we propose UniCustom, a unified visual conditioning framework that fuses ViT and VAE features before VLM encoding. This early fusion exposes the VLM to both semantic cues and appearance-rich details, enabling its hidden states to jointly encode the referred subject and corresponding visual appearance with only a lightweight linear fusion layer. To learn such unified representations, we adopt a two-stage training strategy: reconstruction-oriented pretraining that preserves reference-specific appearance details in the fused hidden states, followed by supervised finetuning on single- and multi-reference generation tasks. We further introduce a slot-wise binding regularization that encourages each image slot to preserve low-level details of its corresponding reference, thereby reducing cross-reference entanglement. Experiments on two multi-reference generation benchmarks demonstrate that UniCustom consistently improves subject consistency, instruction following, and compositional fidelity over strong baselines.
Yiyan Xu, Qiulin Wang, Wenjie Wang +5
University of Science and Technology of China · Kling Team, Kuaishou Technology