cs.CVSep 30, 2026

FAST: Flow Any Scene Transformer

Authors: Yongjian Zhang, Longguang Wang, Zhuo Song, Zhiheng Fu, Liang Lin, Yulan Guo

Organizations: School of Electronics and Communication Engineering, the Shenzhen Campus of Sun Yat-sen University, Sun Yat-sen University, Shenzhen, China · Hong Kong Polytechnic University, HKSAR, China · School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China

Abstract

Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SAMatcher: Co-Visibility Modeling with Segment Anything for Robust Feature Matching

    Jun 2, 2026Xu Pan, Qiyuan Ma, Mingyue Dong +3Harder Better Faster Denser Feature MatchingSegment Anything Model

  2. Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

    Aug 11, 2026Songlin Du, Xiaoyong Lu, Zeyu Wu +5Harder Better Faster Denser Feature MatchingCross-View

  3. RoMa v2: Harder Better Faster Denser Feature Matching

    Nov 19, 2025Johan Edstedt, David Nordström, Yushan Zhang +7Harder Better Faster Denser Feature MatchingDense Prediction