RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing
Organizations: Dartmouth College
Abstract
Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive as the number and resolution of references grow. We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms. Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues. Condition routing and attention routing align reference tokens with their assigned target regions and restrict cross-reference interactions, while allowing selective reference access beyond region boundaries for scene integration. We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references. After many-reference fine-tuning, RefRoute achieves an overall Weighted-Ref-VIEScore of 36.06 on ManyRef100, compared with 8.88 for FLUX.2-Klein-9B. Separate inference-cost evaluations show substantially slower latency growth as the reference count increases: at 16 references, our 50-step and 4-step configurations achieve and speedups over their corresponding FLUX baselines, respectively. These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.
Figures & tables
| Method | Object | Person | HOI | De&Re | Overall |
|---|---|---|---|---|---|
| Gemini-3-Pro-Image-Preview | 50.59 | 54.75 | 50.21 | 52.13 | 51.76 |
| GPT-Image-1.5 | 56.66 | 46.16 | 52.35 | 48.46 | 50.60 |
| Gemini-2.5-Flash-Image | 48.01 | 41.79 | 49.64 | 49.44 | 47.83 |
| Qwen-Image-MICo | 52.38 | 21.11 | 34.95 | 37.42 | 35.86 |
| BAGEL-MICo | 38.98 | 28.45 | 25.30 | 44.51 | 34.41 |
| OmniGen2-MICo | 46.26 | 22.85 | 32.18 | 36.82 | 33.82 |
| Method | Overall | Human | Object | Mixed | |||
|---|---|---|---|---|---|---|---|
| FLUX.2-Klein-Base-9B | 4.18 | 10.79 | 0.09 | 2.28 | 0.268 | 1.470 | 6.949 |
| FLUX.2-Klein-9B | 8.88 | 17.07 | 4.56 | 5.98 | 0.375 | 2.679 | 7.319 |
| GPT-Image-1 | 27.92 | 10.31 | 55.95 | 20.09 | 0.574 | 5.613 | 7.763 |
| Ours-MICo | 17.32 | 13.60 | 36.10 | 6.02 | 0.639 | 3.758 | 6.171 |
| Ours-ManyRef | 36.06 | 28.10 | 57.37 | 26.04 | 0.709 | 6.428 | 7.583 |
| (a) Reconstruction | (b) ManyRef100 |
|---|---|
| Design PSNR SSIM DISTS HF-MAE VAE latent DiT 24.22 0.725 0.0837 0.0220 Pixel VAE 23.57 0.724 0.0949 0.0226 Pixel DiT 24.78 0.745 0.0789 0.0207 | Design Human Mixed Object Concat 0.264 / 0.177 / 0.390 0.217 / 0.177 / 0.385 0.701 / 0.426 / 0.794 Pixel VAE 0.254 / 0.168 / 0.413 0.223 / 0.198 / 0.427 0.696 / 0.409 / 0.799 Pixel DiT 0.300 / 0.217 / 0.517 0.248 / 0.221 / 0.467 0.731 / 0.444 / 0.824 |
| Count MAE | |||
|---|---|---|---|
| Setting | Overall | Human | Mix |
| w/o Stage 1 | 1.643 | 1.300 | 1.900 |
| w/ Stage 1 | 1.057 | 1.033 | 1.075 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Group | Models and use |
|---|---|
| Ours | Ours-MICo (40k); Ours-ManyRef (high- continuation of Ours-MICo ). Evaluated by us on MICo-Bench and ManyRef100 at . |
| MICo-tuned open models | reported MICo-Bench scores. |
| Native open models | FLUX.2-Klein-9B (evaluated by us); OmniGen2 and Qwen-Image-Edit (reported MICo-Bench scores). |
| OminiControl2 mechanism | Released reference-branch code reproduced on our backbone; no updated public weights. 4- and 50-step cost only. |
| Stage 1 (T2I) | Stage 2 (joint training) | Stage 3 (distillation) | |
|---|---|---|---|
| Initialized from | Backbone weights | Stage 1 compressor + context LoRA | Trained teacher |
| Objective | (text-to-image) | (multi-ref, routed) | Distillation following TBSM [ Sun et al., 2026 ] |
| Trainable | Compressor, context LoRA | Compressor, context + spatial LoRA, layout stream | LoRA adapters |
| LoRA rank / LR | 128/ | 128/ | 128/ |
| Training steps | 40k | 40k (main); 60k total (high- ) | 20k |
| Batch / GPUs | / 4 | / 8 | / 8 |