ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding
Organizations: Pukyong National University, Busan, Republic of Korea
Abstract
The strong performance that modern semantic correspondence methods achieve at standard thresholds plateaus sharply at fine-grained thresholds. We argue that this plateau stems not from the representational capacity of backbone features, but from a grid-tied readout. Patch-based vision transformers tokenize images onto discrete grids, introducing two forms of quantization error: querying nearest patch features instead of exact keypoints on the source side, and the absence of grid features representing precise ground-truth locations on the target side. We quantify this quantization ceiling across all 499,188 keypoints in SPair-71k: under the standard 448x448, patch-14 setting, 84.9% of ground-truth keypoints have no grid feature representing their precise location at PCK@0.01. This is a structural limitation at the representation level, independent of the matching strategy. We address this with ImCorr: Sub-pixel Semantic Correspondence via Implicit Feature Decoding, which formulates correspondence estimation over a continuous feature field queryable at arbitrary continuous coordinates. A FiLM-conditioned decoder is trained to embed sub-pixel positional information into the feature field. Querying the field directly at exact keypoint coordinates theoretically eliminates representation-level quantization error on the source side, while decoding onto a grid denser than the backbone grid substantially reduces quantization error on the target side. On SPair-71k and AP-10K (intra-species, cross-species, and cross-family), ImCorr improves performance at fine-grained thresholds (PCK@0.01-0.05), achieving a 6.2 percentage point gain over the prior state of the art at PCK@0.01 on SPair-71k. These results demonstrate that representational continuity is an effective solution for precise semantic correspondence. Code is available at https://github.com/YusungChoi/ImCorr.
Figures & tables
| Image Size | Patch Size | PCK@0.1 | PCK@0.05 | PCK@0.01 |
|---|---|---|---|---|
| 224 224 | 14 | 2.70% | 27.56% | 96.23% |
| 16 | 5.30% | 36.72% | 97.22% | |
| 448 448 | 14 | 0.26% | 2.88% | 84.91% |
| 16 | 0.28% | 5.41% | 88.59% | |
| 896 896 | 14 | <0.01% | 0.22% | 44.06% |
| 16 | 0.03% | 0.34% | 55.16% |
| SPair-71k | AP-10K (I.S.) | AP-10K (C.S.) | AP-10K (C.F.) | ||||||||||
| Method | Backbone | 0.01 | 0.05 | 0.10 | 0.01 | 0.05 | 0.10 | 0.01 | 0.05 | 0.10 | 0.01 | 0.05 | 0.10 |
| Unsupervised/Weakly Supervised | |||||||||||||
| DINOv2+NN [ 20 ] | ViT-B | 6.3 | 38.4 | 53.9 | 6.4 | 41.0 | 60.9 | 5.3 | 37.0 | 57.3 | 4.4 | 29.4 | 47.4 |
| DIFT [ 17 ] | SD | 7.2 | 39.7 | 52.9 | 6.2 | 34.8 | 50.3 | 5.1 | 30.8 | 46.0 | 3.7 | 22.4 | 35.0 |
| SD+DINO [ 20 ] | SD+ViT-B | 7.9 | 44.7 | 59.9 | 7.6 | 43.5 | 62.9 | 6.4 | 39.7 | 59.3 | 5.2 | 30.8 | 48.3 |
| DIY-SC [ 4 ] | SD+ViT-B | 10.1 | 53.8 | 71.6 | - | - | 70.6 | - | - | 69.8 | - | - | 57.8 |
| Setting | Src | Tgt | PCK@0.01 | PCK@0.05 | PCK@0.10 |
|---|---|---|---|---|---|
| Baseline | 19.1 | 72.9 | 83.3 | ||
| + Src Query | ✓ | 25.6 | 77.2 | 86.7 | |
| + Tgt Decoding | ✓ | 30.1 | 76.1 | 84.1 | |
| Full model (Ours) | ✓ | ✓ | 33.2 | 78.9 | 86.2 |
| Gain over baseline | +14.1 | +6.0 | +2.9 |
| PCK@0.01 | PCK@0.05 | PCK@0.10 | |
|---|---|---|---|
| 2 | 30.3 | 76.2 | 86.9 |
| 4 | 33.2 | 78.9 | 86.2 |
| 6 | 32.9 | 78.0 | 85.6 |
| 8 | 31.7 | 77.1 | 85.1 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.