cs.CVOct 6, 2026

RefRoute: Decoupling Conditioning Cost from References via Compact Residual Conditioning and Spatial Routing

Authors: Wanning He, Yuyao Zhang, Yu-Wing Tai

Organizations: Dartmouth College

Abstract

Multi-reference image generation requires preserving the appearance of multiple subjects while composing them into a coherent scene. However, existing diffusion transformers commonly encode references as dense visual token grids and jointly process them with global attention, making conditioning increasingly expensive as the number and resolution of references grow. We present RefRoute, a framework that addresses both reference representation cost and attention overhead through two complementary mechanisms. Compact residual conditioning combines low-resolution latent tokens with lightweight residual features extracted from full-resolution pixels, reducing reference token counts while retaining fine-grained appearance cues. Condition routing and attention routing align reference tokens with their assigned target regions and restrict cross-reference interactions, while allowing selective reference access beyond region boundaries for scene integration. We further introduce RefRoute-Data for training many-reference generation models and ManyRef100, a benchmark spanning human, object, and mixed compositions with 10-17 references. After many-reference fine-tuning, RefRoute achieves an overall Weighted-Ref-VIEScore of 36.06 on ManyRef100, compared with 8.88 for FLUX.2-Klein-9B. Separate inference-cost evaluations show substantially slower latency growth as the reference count increases: at 16 references, our 50-step and 4-step configurations achieve 18.3×18.3\times and 14.2×14.2\times speedups over their corresponding FLUX baselines, respectively. These results establish compact reference representations and spatially routed attention as an effective approach to scalable many-reference image generation.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping

    Jun 22, 2026Rishubh Parihar, Ayush Raina, R. Venkatesh Babu +1Diffusion ModelsHigh-Fidelity Conditional Generation

  2. Beyond Attention Masks: Instruction Anchoring for Efficient In-Context Diffusion Generation

    Aug 21, 2026Yangshuai Liu, Zheming Li, Jiaao Li +4Diffusion TransformersMultimodal Generation

  3. UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation

    May 12, 2026Yiyan Xu, Qiulin Wang, Wenjie Wang +5Multi-Reference Image GenerationDiffusion Models