UltraMatch: Transport Path Routing for Ultra-Fast and Memory-Efficient Image Matching
Organizations: Electronic Information School, Wuhan University, China · Xiaomi Corporation, China · School of Computer Science and Technology, Harbin Institute of Technology, China · School of Robotics, Wuhan University, China
Abstract
Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framework that bypasses the quadratic computation and memory cost of dense token-level matching by routing only a small fraction of candidate matching paths. At its core, a lightweight Transport Path Router operates on coarse block representations to rank candidate target blocks for each source block and retain only a small set, restricting subsequent token-level matching to the selected paths and avoiding the construction of the full token-to-token matching matrix. We further design a sparse global Dual-Softmax that performs matching only over the routed block candidates while retaining global competition across the sparse matching space. Beyond matching acceleration, UltraMatch employs deployment-oriented structural reparameterization for feature extraction and a tiny fine matching head with shared parameters, further reducing inference cost and memory consumption. UltraMatch achieves competitive accuracy among semi-dense matchers, while running 1.67 faster than SuperPoint+LightGlue with only 0.44 GiB peak inference memory. Its scalability enables inference at up to 6K resolution on a single RTX 3090, whereas existing semi-dense matchers run out of memory before reaching 2K. Our routing strategy is also transferable, delivering about 2 end-to-end speedup in EDM and ELoFTR without accuracy loss. The project repository is available at https://github.com/JiajunLe/UltraMatch.
Figures & tables
| Category | Method | ScanNet-1500 ( ) | MegaDepth-1500 ( ) | Time (ms)( ) | Memory (GiB)( ) | ||||
| AUC@ | AUC@ | AUC@ | AUC@ | AUC@ | AUC@ | ||||
| Sparse | SP + SG | 16.2 | 32.8 | 49.7 | 49.7 | 67.1 | 80.6 | 104.83 | 1.02 |
| SP + LG | 14.8 | 30.8 | 47.5 | 49.9 | 67.0 | 80.1 | 53.96 | 1.02 | |
| Dense | DKM | 26.6 | 47.1 | 64.2 | 60.4 | 74.9 | 85.1 | 572.98 | 9.66 |
| RoMa | 28.9 | 50.4 | 68.3 | 62.6 | 76.7 | 86.3 | 755.44 | 6.79 | |
| Semi-Dense | LoFTR | 16.9 | 33.6 | 50.6 | 52.8 | 69.2 | 81.2 | 376.15 | 12.19 |
| Method | HPatches ( ) | Aachen Day-Night v1.1 ( ) | InLoc ( ) | ||||
| @3px | @5px | @10px | Day | Night | DUC1 | DUC2 | |
| ( ) / ( ) / ( ) | ( ) / ( ) / ( ) | ||||||
| SP + SG | 37.0 | 52.6 | 70.1 | 89.7 / 96.5 / 99.3 | 73.8 / 91.1 / 99.5 | 50.0 / 69.7 / 79.8 | 47.3 / 77.9 / 80.2 |
| SP + LG | 35.7 | 51.6 | 70.2 | 89.2 / 96.5 / 99.3 | 72.3 / 89.5 / 99.0 | 48.0 / 68.7 / 79.8 | 44.3 / 71.0 / 75.6 |
| LoFTR | 51.3 | 63.0 | 75.5 | 88.7 / 96.1 / 98.6 | 77.0 / 90.6 / 99.5 | 49.0 / 71.7 / 84.3 | 51.1 / 73.3 / 81.7 |
| MatchFormer | 51.4 | 63.1 | 76.1 | 89.4 / 96.0 / 98.8 | 75.9 / 90.6 / 99.5 | 50.0 / 73.7 / 85.4 | 58.0 / 80.9 / 87.0 |
| AUC@5 ∘ | T. (ms) | Mem. (GiB) | Rec. (%) | |
| 1 | 54.5 | 31.53 | 0.44 | 82.2 |
| 4 | 55.6 | 31.63 | 0.44 | 97.9 |
| 6 | 57.4 | 32.26 | 0.44 | 98.8 |
| 8 | 56.6 | 33.87 | 0.45 | 99.2 |
| Dense | 56.5 | 74.94 | 8.26 | 100.0 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| Resolution level | Input size ( ) |
| 0.5K | |
| 0.75K | |
| 1K | |
| 1.25K | |
| 1.5K | |
| 1.8K |
| Stage | Time (ms) | Ratio (%) |
| Feature Extraction | 10.80 | 33.48 |
| Feature Interaction | 12.13 | 37.60 |
| Router | 0.32 | 0.99 |
| Coarse Matching | 4.27 | 13.24 |
| Refinement | 4.17 | 12.93 |
| Other Overhead | 0.57 | 1.77 |
| @5 ∘ | @10 ∘ | T. (ms) | |||
| 52.4 | 68.5 | 32.12 | |||
| ✓ | 55.7 | 72.0 | 32.88 | ||
| ✓ | ✓ | 56.6 | 72.5 | 32.48 | |
| ✓ | ✓ | ✓ | 57.4 | 72.5 | 32.26 |
| Variant | AUC@5 ∘ | AUC@10 ∘ | Time (ms) | Memory (GiB) |
| Dense Dual-Softmax (Native) | 56.5 | 72.4 | 74.94 | 8.26 |
| Dense Dual-Softmax (Streaming) | 56.5 | 72.4 | 289.22 | 0.43 |
| Transport Path Routing | 57.4 | 72.5 | 32.26 | 0.44 |
| Method | Input | Configuration | Precision | AUC@5 ∘ | # Matches | Time (ms) | Memory (GiB) |
| SuperPoint+LightGlue | Adaptive, † | FP32 + FP16 attn. † | 49.7 | 553 | 53.96 | 1.02 | |
| ELoFTR-Full | Full, | Mixed FP16 | 56.4 | 3288 | 140.49 | 8.64 | |
| ELoFTR-Opt | Opt, | Mixed FP16 | 55.4 | 3531 | 92.17 | 3.43 | |
| EDM | FP32 | 57.5 | 4326 | 83.45 | 8.24 | ||
| UltraMatch | , halo , | FP32 | 57.4 | 3973 | 37.85 | 0.57 | |
| UltraMatch | , halo , | Selective BF16 | 57.4 | 3973 | 32.26 | 0.44 |