UltraMatch: Transport Path Routing for Ultra-Fast and Memory-Efficient Image Matching
Authors: Jiajun Le, Yifan Lu, Zizhuo Li, Lei Cao, Junjun Jiang, Jiayi Ma
Organizations: Electronic Information School, Wuhan University, China · Xiaomi Corporation, China · School of Computer Science and Technology, Harbin Institute of Technology, China · School of Robotics, Wuhan University, China
Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framework that bypasses the quadratic computation and memory cost of dense token-level matching by routing only a small fraction of candidate matching paths. At its core, a lightweight Transport Path Router operates on coarse block representations to rank candidate target blocks for each source block and retain only a small set, restricting subsequent token-level matching to the selected paths and avoiding the construction of the full token-to-token matching matrix. We further design a sparse global Dual-Softmax that performs matching only over the routed block candidates while retaining global competition across the sparse matching space. Beyond matching acceleration, UltraMatch employs deployment-oriented structural reparameterization for feature extraction and a tiny fine matching head with shared parameters, further reducing inference cost and memory consumption. UltraMatch achieves competitive accuracy among semi-dense matchers, while running 1.67× faster than SuperPoint+LightGlue with only 0.44 GiB peak inference memory. Its scalability enables inference at up to 6K resolution on a single RTX 3090, whereas existing semi-dense matchers run out of memory before reaching 2K. Our routing strategy is also transferable, delivering about 2× end-to-end speedup in EDM and ELoFTR without accuracy loss. The project repository is available at https://github.com/JiajunLe/UltraMatch.
Figures & tables
Figure 1: Scalability and efficiency of UltraMatch. (a) Runtime comparison across increasing input resolutions. (b) Efficiency trade-off in latency and memory consumption on MegaDepth-1500, where bubble size indicates AUC@5 ∘ . All measurements are conducted on a single RTX 3090 GPU.
Figure 2: Overview of UltraMatch. (a) A lightweight backbone extracts multi-scale features, with reparameterizable blocks fused into standard convolutions at inference. (b) The features F32 undergo iterative self- and cross-attention, and are propagated to 1/8 through RepFuse. (c) The interacted F32 route candidate block pairs, whose corresponding 1/8 features are indexed for sparse token matching. The resulting scores are assembled into a sparse global matrix S and normalized by Sparse Global Dual-Softmax to obtain coarse correspondences Mc . (d) For each coarse correspondence, indexed Ff and F8 are fused and processed by shared lightweight encoders, followed by axis-wise heads that predict offset distributions and uncertainties for subpixel refinement.
Figure 3: Efficiency and scalability comparison under increasing input resolutions. Inference runtime, peak inference memory, and peak training memory are reported. The horizontal dashed line indicates the effective memory capacity of an NVIDIA RTX 3090, while dashed curve extensions terminated by crosses denote out-of-memory (OOM) points.
Category
Method
ScanNet-1500 ( ↑ )
MegaDepth-1500 ( ↑ )
Time (ms)( ↓ )
Memory (GiB)( ↓ )
AUC@ 5∘
AUC@ 10∘
AUC@ 20∘
AUC@ 5∘
AUC@ 10∘
AUC@ 20∘
Sparse
SP + SG
16.2
32.8
49.7
49.7
67.1
80.6
104.83
1.02
SP + LG
14.8
30.8
47.5
49.9
67.0
80.1
53.96
1.02
Dense
DKM
26.6
47.1
64.2
60.4
74.9
85.1
572.98
9.66
RoMa
28.9
50.4
68.3
62.6
76.7
86.3
755.44
6.79
Semi-Dense
LoFTR
16.9
33.6
50.6
52.8
69.2
81.2
376.15
12.19
Table 1: Relative pose estimation on ScanNet and MegaDepth. Pose AUCs at different thresholds are reported alongside runtime and peak GPU memory per image pair on MegaDepth. Bold and underlined values indicate the best and second-best results among semi-dense methods, respectively.
Method
HPatches ( ↑ )
Aachen Day-Night v1.1 ( ↑ )
InLoc ( ↑ )
@3px
@5px
@10px
Day
Night
DUC1
DUC2
( 0.25m,2∘ ) / ( 0.5m,5∘ ) / ( 5.0m,10∘ )
( 0.25m,2∘ ) / ( 0.5m,5∘ ) / ( 1.0m,10∘ )
SP + SG
37.0
52.6
70.1
89.7 / 96.5 / 99.3
73.8 / 91.1 / 99.5
50.0 / 69.7 / 79.8
47.3 / 77.9 / 80.2
SP + LG
35.7
51.6
70.2
89.2 / 96.5 / 99.3
72.3 / 89.5 / 99.0
48.0 / 68.7 / 79.8
44.3 / 71.0 / 75.6
LoFTR
51.3
63.0
75.5
88.7 / 96.1 / 98.6
77.0 / 90.6 / 99.5
49.0 / 71.7 / 84.3
51.1 / 73.3 / 81.7
MatchFormer
51.4
63.1
76.1
89.4 / 96.0 / 98.8
75.9 / 90.6 / 99.5
50.0 / 73.7 / 85.4
58.0 / 80.9 / 87.0
Table 2: Evaluation of homography estimation on HPatches and visual localization on Aachen Day-Night v1.1 and InLoc. Homography AUCs at 3, 5, and 10 pixels and percentages of correctly localized queries at three thresholds are reported.
Table 6
R
AUC@5 ∘
T. (ms)
Mem. (GiB)
Rec. (%)
1
54.5
31.53
0.44
82.2
4
55.6
31.63
0.44
97.9
6
57.4
32.26
0.44
98.8
8
56.6
33.87
0.45
99.2
Dense
56.5
74.94
8.26
100.0
Table 5: Routing Size Analysis.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Resolution level
Input size ( W×H )
0.5K
480×320
0.75K
768×512
1K
960×640
1.25K
1248×832
1.5K
1536×1024
1.8K
1824×1216
Appendix
Table 6: Input-resolution settings used in the high-resolution experiments. Each resolution denotes the size of an individual input image.
Figure 4: High-resolution matching analysis on ETH3D. (a) Relative pose AUC@5 ∘ under increasing input resolutions. UltraMatch maintains strong geometric accuracy up to 6K resolution, whereas competing semi-dense matchers either reach their memory limits at substantially lower resolutions or suffer pronounced accuracy degradation. (b) Ground-truth route coverage, comparing the exact Top-6 routed blocks with the candidate space after halo expansion. Although GT coverage decreases as the token space grows, halo expansion consistently recovers additional valid correspondence paths. (c) Proportion of total inference time spent on routing score computation, which remains below 3% across all evaluated resolutions.
Figure 5: Qualitative matching comparisons of UltraMatch, LoFTR, and ELoFTR on MegaDepth and ScanNet. Green lines indicate geometrically consistent matches, while red lines denote matches whose epipolar error exceeds 5×10−4 in normalized image coordinates.
Stage
Time (ms)
Ratio (%)
Feature Extraction
10.80
33.48
Feature Interaction
12.13
37.60
Router
0.32
0.99
Coarse Matching
4.27
13.24
Refinement
4.17
12.93
Other Overhead
0.57
1.77
Appendix
Table 7: Runtime breakdown of UltraMatch.
Lr
Lcov
Lrank
@5 ∘
@10 ∘
T. (ms)
52.4
68.5
32.12
✓
55.7
72.0
32.88
✓
✓
56.6
72.5
32.48
✓
✓
✓
57.4
72.5
32.26
Appendix
Table 8: Ablation study of routing losses on MegaDepth.
Variant
AUC@5 ∘ ↑
AUC@10 ∘ ↑
Time (ms) ↓
Memory (GiB) ↓
Dense Dual-Softmax (Native)
56.5
72.4
74.94
8.26
Dense Dual-Softmax (Streaming)
56.5
72.4
289.22
0.43
Transport Path Routing
57.4
72.5
32.26
0.44
Appendix
Table 9: Analysis of matching sparsification and memory-efficient normalization on MegaDepth. Relative pose AUC (%), average pairwise matching time, and peak allocated GPU memory at 1152×1152 resolution are reported.
Method
Input
Configuration
Precision
AUC@5 ∘
# Matches
Time (ms)
Memory (GiB)
SuperPoint+LightGlue
11522
Adaptive, K≤2048 †
FP32 + FP16 attn. †
49.7
553
53.96
1.02
ELoFTR-Full
11522
Full, θc=0.1
Mixed FP16
56.4
3288
140.49
8.64
ELoFTR-Opt
11522
Opt, θc=20
Mixed FP16
55.4
3531
92.17
3.43
EDM
11522
θc=0.05
FP32
57.5
4326
83.45
8.24
UltraMatch
11522
R=6 , halo =1 , θc=0.10
FP32
57.4
3973
37.85
0.57
UltraMatch
11522
R=6 , halo =1 , θc=0.10
Selective BF16
57.4
3973
32.26
0.44
Appendix
Table 10: Controlled efficiency comparison and inference configurations on MegaDepth-1500. Accuracy, correspondence count, runtime, and memory are measured under the listed inference configuration. † “Adaptive, K≤2048 ” denotes at most 2048 keypoints per image, with early stopping and point pruning enabled. “FP32 + FP16 attn.” denotes FP32 inference with Q/K/V internally cast to FP16 in the Flash attention path.
Dense feature matching aims to estimate all correspondences between two images of a 3D scene and has recently been established as the gold standard due to its high accuracy and robustness. However, existing dense matchers still fail or perform poorly for many hard real-world scenarios, and high-precision models are often slow, limiting their applicability. In this paper, we attack these weaknesses on a wide front through a series of systematic improvements that together yield a significantly better model. In particular, we construct a novel matching architecture and loss, which, combined with a curated diverse training distribution, enables our model to solve many complex matching tasks. We further make training faster through a decoupled two-stage matching-then-refinement pipeline, and at the same time, significantly reduce refinement memory usage through a custom CUDA kernel. Finally, we leverage the recent DINOv3 foundation model along with multiple other insights to make the model more robust and unbiased. In our extensive set of experiments, we show that the resulting novel matcher sets a new state-of-the-art, being significantly more accurate than its predecessors. Code is available at https://github.com/Parskatt/romav2
Johan Edstedt, David Nordström, Yushan Zhang +7
Linköping University · Chalmers University of Technology · University of Amsterdam +1
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs. Consequently, current frameworks typically freeze VFMs and shift the entire burden of rotation generalization to the downstream decoder. To break this architectural bottleneck, we propose REDI-Match, an efficient framework driven by a novel Rotation-Equivariant Distillation (REDI) paradigm. Instead of relying on rotation data augmentation to establish rotational correspondences, REDI distills the non-equivariant semantic representations of a VFM into a lightweight, strictly rotation-equivariant encoder, leveraging an equivariant geometric architecture to constrain robust high-dimensional semantics. To fully exploit these features, we equip the decoder with an entropy-driven spatial alignment module. By evaluating discrete rotation hypotheses, this mechanism explicitly locks onto the canonical coordinate system, eliminating global ambiguity before continuous refinement. Extensive experiments demonstrate that REDI-Match establishes a new state-of-the-art (SOTA) across multiple benchmarks. Notably, it achieves a 13.89% absolute pose accuracy improvement on the highly challenging SatAst dataset while operating 1.9x faster than the current SOTA (RoMa v2), enabling real-time inference (~41 FPS) on a single RTX 4090 GPU. Code: https://github.com/YinjiGe/REDI-Match.
Yinji Ge, Guixu Zheng, Wulong Guo +5
1Tsinghua University · 2Southern University of Science and Technology · 3Beihang University +1
Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains underexplored. In this work, we present Flow Any Scene Transformer (FAST), a scalable correspondence model driven by two key insights. First, we reveal that the query-key projections inside single-view vision foundation models encode a coarse yet reusable prior for cross-view matching. Second, reusing these pretrained projections in cross-attention form yields a highly effective initialization for a ViT-based matcher built from a single-view encoder. Guided by these insights, we build FAST upon a vanilla single-view foundation model, utilizing a zero-parameter rewiring strategy to convert selected self-attention layers into cross-attention for cross-view interaction. This design allows ViT-based matchers to scale with advances in single-view foundation models, bypassing the need for a dedicated pair-centric pretraining stage. To fully unlock the scaling potential of this formulation, we assemble a 6-million-pair training corpus for general-purpose dense 2D displacement estimation across diverse co-visible image pairs. Extensive experiments demonstrate that FAST achieves state-of-the-art performance across a wide range of benchmarks, while scaling favorably with both backbone size and training data.
Yongjian Zhang, Longguang Wang, Zhuo Song +3
School of Electronics and Communication Engineering, the Shenzhen Campus of Sun Yat-sen University, Sun Yat-sen University, Shenzhen, China · Hong Kong Polytechnic University, HKSAR, China · School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China