FAST: Flow Any Scene Transformer
Organizations: School of Electronics and Communication Engineering, the Shenzhen Campus of Sun Yat-sen University, Sun Yat-sen University, Shenzhen, China · Hong Kong Polytechnic University, HKSAR, China · School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China
Abstract
Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.
Figures & tables
| Method | Sintel | Spring | KITTI | |||||
|---|---|---|---|---|---|---|---|---|
| clean | final | EPE | 1PE | F1 | F1-all | F1-bg | F1-fg | |
| PWCNet [ 11 ] | 3.45 | 4.60 | 2.288 | 82.27 | 4.889 | 7.72 | 7.69 | 7.88 |
| RAFT [ 12 ] | 1.61 | 2.86 | 1.476 | 6.790 | 3.198 | 5.10 | 4.74 | 6.87 |
| SEA-RAFT (M) [ 3 ] | 1.44 | 2.86 | 0.363 | 3.686 | 1.347 | 4.64 | 4.47 | 5.49 |
| CrocoFlow [ 23 ] | 1.09 | 2.44 | 0.498 | 4.565 | 1.508 | 3.64 | 3.18 | 5.94 |
| FlowFormer++ [ 42 ] | 1.07 | 1.94 | - | - | - | 4.52 | - | - |
| Methods | KT15 | KT12 | ETH3D | Midd. |
|---|---|---|---|---|
| D1-all | D1-all | 1PE-noc | 2PE-noc | |
| Selective-IGEV [ 43 ] | 4.5 | 3.2 | 3.4 | 7.5 |
| MatchAttention-B [ 44 ] | 3.78 | - | 1.27 | 2.27 |
| S2M2-XL [ 41 ] | 2.97 | 4.05 | 0.42 | 1.0 |
| FoundationStereo [ 7 ] | 2.8 | 2.3 | 0.5 | 1.1 |
| FlowFormer++ [ 42 ] | 6.12 | 4.50 | 4.63 | 10.99 |
| Method | DTU | TA-WB | ||||||
|---|---|---|---|---|---|---|---|---|
| epe | 1pe | 2pe | 5pe | epe | 1pe | 2pe | 5pe | |
| PanMatch [ 9 ] | 26.69 | 60.7 | 44.8 | 32.7 | 74.36 | 69.7 | 55.8 | 48.0 |
| UFM [ 25 ] | 5.55 | 55.5 | 32.9 | 13.8 | 12.84 | 51.4 | 30.6 | 17.0 |
| RoMav2 [ 27 ] | 4.81 | 38.8 | 20.5 | 9.7 | 10.73 | 38.9 | 18.6 | 11.5 |
| FAST | 4.43 | 54.0 | 28.7 | 10.4 | 12.03 | 51.0 | 28.6 | 14.7 |
| Method | WxBS | ScanNet | MegaDepth | ||||
|---|---|---|---|---|---|---|---|
| @10px | @5 ∘ | @10 ∘ | @20 ∘ | @5 ∘ | @10 ∘ | @20 ∘ | |
| LightGlue [ 55 ] | – | 17.8 | 34.0 | 52.0 | 51.0 | 68.1 | 80.7 |
| LoFTR [ 54 ] | 55.4 | 22.1 | 40.8 | 57.6 | 52.8 | 69.2 | 81.2 |
| DKM [ 56 ] | 58.9 | 29.4 | 50.7 | 68.3 | 60.4 | 74.9 | 85.1 |
| RoMa [ 8 ] | 80.1 | 31.8 | 53.4 | 70.9 | 62.6 | 76.7 | 86.3 |
| UFM [ 25 ] | 42.3 | 31.3 | 54.1 | 72.0 | 41.5 | 57.9 | 72.4 |
| Experiment | Stereo Matching | Optical Flow | ||||||
|---|---|---|---|---|---|---|---|---|
| Middlebury | ETH3D | Sintel | KITTI 15 | |||||
| EPE | 2PE | EPE | 1PE | clean | final | EPE | F1-all | |
| Scratch | 5.94 | 43.47 | 0.79 | 16.70 | 3.02 | 4.11 | 4.84 | 21.52 |
| DINOv2 [ 26 ] | 2.01 | 16.84 | 0.37 | 5.02 | 1.21 | 2.28 | 2.27 | 8.85 |
| DINOv3 [ 6 ] | 1.77 | 13.18 | 0.28 | 3.10 | 1.06 | 2.13 | 2.06 | 7.33 |
| global SA | 2.38 | 17.52 | 0.31 | 4.32 | 1.42 | 2.44 | 2.74 | 11.52 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Backbone | Cross-view Interaction | Params. | Decoder |
|---|---|---|---|---|
| CroCo v2 [ 23 ] | ViT [ 57 ] | ViT-B | 114M | DPT |
| UFM [ 25 ] | DINOv2 [ 26 ] | 12 Transformer | 86M | DPT |
| RoMa v2 [ 27 ] | DINOv3 [ 6 ] | 12 Transformer + CNN | 90M | DPT + Linear |
| FAST (Ours) | DINOv3 [ 6 ] | Rewired Transformer | 0 | DPT |
| Dataset | Dynamic Scenes | Lables | Res. | (Seq, Stride) | Clips | ||
| Flow | Depth | Camera | |||||
| AutoFlow [ 58 ] | ✓ | ✓ | ✗ | ✗ | 448 576 | (2,1) | 27295 |
| cvo [ 59 ] | ✓ | ✓ | ✗ | ✗ | 512 512 | (2,1) | 125466 |
| DynamicReplica [ 60 ] | ✓ | ✓ | ✓ | ✓ | 720 1280 | (8,4) | 35742 |
| FlyingChairs [ 10 ] | ✓ | ✓ | ✗ | ✗ | 384 512 | (2,1) | 22872 |
| Infinigen [ 61 ] | ✓ | ✓ | ✓ | ✓ | 720 1280 | (2,1) | 1910 |