Image Matching

Latest papers 52

Oct 7, 2026cs.CV

An Invariant Tangent-Angle Descriptor and a Band U-Net for 2D Fragment Adjacency Prediction

This paper addresses the prediction of adjacency between pairs of 2D fragments based on their contours. We improved the two-stage architecture proposed in Beaulac's thesis, in which a rotation-equivariant Siamese convolutional neural network scores pairs of local image windows along the two contours of two fragments. The scores are gathered in an adjacency matrix in which a ResNet detects the partial anti-diagonal band that reveals the adjacency of two fragments. In the current work, we keep the pipeline and replace the local score by a comparison of tangent-angle profiles of contour windows, making it, by construction, invariant to fragment rotation and agnostic to the selected contour-starting point. These adaptations may be either a training-free likelihood ratio or a small one-dimensional convolutional model trained on corresponding points. We also replaced the final classifier by a band U-Net that segments the band and classifies the pair, so that the shared arc is obtained along with the decision. In the synthetic data set of the original thesis, the tangent descriptor performs as well as or better than the image-window approach in all tested configurations. The proposed pipeline reaches an accuracy of 98%, vs 93% to 95% for the original approach once its evaluation is corrected. We tested our pipeline, with models trained only on synthetic data, on the PairingNet benchmark, and obtained an AUC of 0.93. Furthermore, under the PairingNet pair-searching protocol conditions, our learned descriptor obtains a Recall@10 of 0.82 on the real set against 0.56 from the best model of the original paper.
Oct 5, 2026cs.CV

WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification

Individual animal re-identification from camera-trap imagery is an instance retrieval problem central to non-invasive wildlife monitoring: a query image must retrieve the correct individual from a reference set of known animals. This requires computer vision models to recognize distinctive local patterns in fur, skin, or other visual markings. Current approaches either learn global embeddings as a classification problem, requiring many labeled images per individual while largely ignoring local evidence, or apply off-the-shelf, domain-agnostic image matchers. Although such matchers are pretrained on large and diverse image collections, adapting them to wildlife imagery is challenging because available datasets are small and lack correspondence-level annotations. We study weakly supervised adaptation of a pretrained keypoint matcher using only identity labels, without keypoint-level or geometric correspondence ground truth. We mine informative image pairs with the pretrained matcher, derive weak positive and negative supervision from identity agreement, and contrastively fine-tune the matching network to strengthen correspondences for same-identity pairs and suppress them for different identities. Across open-source wildlife re-identification datasets, our approach improves accuracy over off-the-shelf matchers and a state-of-the-art local--global fusion method. Under an open-world protocol with held-out individuals, it learns a transferable correspondence prior rather than memorizing training identities. To our knowledge, this is the first study of matcher-level, identity-supervised adaptation for animal re-identification. Our method enables data-efficient specialization of image matching models to wildlife domains using identity annotations already available in typical monitoring datasets.
Sep 30, 2026cs.CV

Front-to-Back: Benchmarking Vision-Language Models for Asymmetric Cross-View Vehicle Re-Identification

Matching the same vehicle across front and rear cameras is difficult because the cameras do not share a view and the vehicle's appearance changes substantially. We introduce Front2Back-ReID, a benchmark of 500 manually verified vehicle handovers from 20 recording sequences in South Africa. Each example asks a model to match a vehicle highlighted in a front-camera image to the same vehicle among at least three candidates in a later rear-camera image. We evaluate seven zero-shot vision-language models, four image-retrieval baselines, and 25 human participants. Models are tested using full front RGB images, cropped target vehicles, and binary silhouettes. The strongest VLM achieved 76.6 percent Rank-1 accuracy on target crops, compared with 74.0 percent for the frozen SigLIP2 baseline; this difference was not statistically clear. Human participants achieved 94.0 percent accuracy with full images and 92.2 percent with target crops. Under our evaluation setup, enabling reasoning improved accuracy across all three input conditions for every model evaluated in both modes. We also found that VLMs generally performed worse on full scenes than on target crops. These results show that general-purpose VLMs do not yet consistently outperform strong visual retrieval for front-to-rear vehicle matching, while humans remain substantially more reliable.
Sep 29, 2026cs.CV

UltraMatch: Transport Path Routing for Ultra-Fast and Memory-Efficient Image Matching

Despite recent advances in accuracy and efficiency, coarse matching remains an indispensable yet costly stage in existing semi-dense matchers due to dense token-level matching. We present UltraMatch, an ultra-efficient and scalable semi-dense matching framework that bypasses the quadratic computation and memory cost of dense token-level matching by routing only a small fraction of candidate matching paths. At its core, a lightweight Transport Path Router operates on coarse block representations to rank candidate target blocks for each source block and retain only a small set, restricting subsequent token-level matching to the selected paths and avoiding the construction of the full token-to-token matching matrix. We further design a sparse global Dual-Softmax that performs matching only over the routed block candidates while retaining global competition across the sparse matching space. Beyond matching acceleration, UltraMatch employs deployment-oriented structural reparameterization for feature extraction and a tiny fine matching head with shared parameters, further reducing inference cost and memory consumption. UltraMatch achieves competitive accuracy among semi-dense matchers, while running 1.67×\times faster than SuperPoint+LightGlue with only 0.44 GiB peak inference memory. Its scalability enables inference at up to 6K resolution on a single RTX 3090, whereas existing semi-dense matchers run out of memory before reaching 2K. Our routing strategy is also transferable, delivering about 2×\times end-to-end speedup in EDM and ELoFTR without accuracy loss. The project repository is available at https://github.com/JiajunLe/UltraMatch.
Sep 29, 2026cs.CV

SCCM: Spherically Consistent Coarse Matching for ERP Dense Feature Correspondence

Equirectangular projection (ERP) is the standard representation for 360∘^\circ imagery, and robust dense feature matching on ERP underpins panoramic stereo, view synthesis, and omnidirectional SLAM. Dense matchers trained on flat images degrade systematically on ERP because the chart introduces three coupled distortions -- topological, metric, and area -- that standard coarse matching and visibility estimation do not explicitly model. We show that correcting the three distortions at the coarse-stage interfaces where they arise -- pairwise distortions in attention, per-pixel distortion in covisibility gating -- improves PCK@1∘1^\circ from 0.229 to 0.275 on Matterport3D under a fixed coarse scaffold, with the refiner architecture unchanged -- our central result. Concretely, SCCM (Spherically Consistent Coarse Matching) augments a chart-naive cross-attention/dual-softmax coarse matcher with two sphere-derived priors: Spherical Positional Attention (SPA) pairs a yaw-periodic RoPE (topology) with a tangent-plane bias (metric), and Area-Aware Covisibility (AAC) applies a pre-sigmoid log-area correction (area). The chart-naive scaffold serves as a controlled reference, separating the scaffold-replacement effect from the spherical-prior effect. Instantiated in the RoMa V1 framework with the same frozen encoder, refiner architecture, and loss, SCCM also outperforms the ERP-native EDM (0.163) and an ERP-retrained RoMa V1 (0.198) under a unified ERP dense matching protocol, while perspective-trained matchers largely fail on ERP. It further transfers zero-shot to Stanford2D3D and, when trained on outdoor Holo360D, leads there as well.
Sep 22, 2026cs.RO

MatcherCompass: A Deployment-Aware Benchmark to Guide Image Matcher Selection in the Wild

Field robots operating across time of day and sensing modalities require accurate image correspondences within onboard time and resource budgets. However, accuracy and runtime reported for individual methods on a single device provide limited guidance for choosing a matcher and its configuration on a target platform. We present MatcherCompass, a deployment-aware benchmark for choosing local feature matchers in field robotics. Under common input and pose-evaluation procedures, we compare nine classical and learned matching pipelines across four image resolutions and supported numerical precisions. Four visual conditions cover viewpoint variation, day--night matching in visible and thermal imagery, and daytime visible--thermal matching. We evaluate pose accuracy using the area under the error--recall curve (AUC) at 5∘5^\circ, 10∘10^\circ, and 20∘20^\circ, and measure runtime, GPU memory, and energy per image pair on four GPU platforms spanning workstation and onboard computers. The results show that changes in hardware, input resolution, and numerical precision can move a matcher across a runtime budget boundary, altering the feasible choices. We organize the measurements into a selection guide that returns all configurations satisfying user-specified time and resource limits, together with their accuracy under the selected visual condition. MatcherCompass provides measured evidence for choosing matching pipelines that fit a robot's sensing conditions and computing hardware. Project page: https://matchercompass.github.io/.
Sep 21, 2026cs.CV

What Survives on Real Drawings: Active Sampling, Connectome Wiring, and Matched Baselines in Architectural Document Vision

A connectome-constrained model of the fly visual system, optimized for motion and then frozen, can be driven over architectural drawings by prescribed motion and used as a texture representation. We compare it with information-matched baselines that see the same 721 photoreceptor samples. On clean synthetic data the frozen model transfers but loses to task training: 0.857 area-weighted accuracy in one-shot hatch matching versus 0.959 for a 5,888-parameter CNN, and 0.619 IoU in wall segmentation versus 0.905 for a matched network. Under scan noise and thickened strokes, the trained networks lose up to 0.188 accuracy while the frozen pipeline loses 0.030. On fourteen production sheets, opened once, a 1,876-parameter fly model reaches 0.505 average precision versus 0.415 for a network two hundred times larger. A preregistered held-out split confirms the clean-data ordering: 0.835 for the circuit, 0.894 for receptors only, and 0.971-0.980 for trained CNNs. Rewiring the connectome while preserving degrees or type pairs and transmitter signs costs 0.271-0.356 accuracy across three seeds, so the exact wiring is load-bearing. Yet the intact circuit does not beat its moving retina, and T4/T5 silencing leaves both tasks intact. Longer observations reverse the circuit-receptor ordering once the stimulus spans a period, but not through T4/T5. Thus active sampling and exact structure matter, while clean-data practical performance remains dominated by task-trained networks and the useful transfer margin is largely retinal.
Sep 21, 2026astro-ph.IM

Scale-Vector Alignment: A Scale-Aware Framework for Spatially Resolved Morphological Similarity in Astronomical Images

Astronomical maps made with different tracers are not expected to have identical morphology. Excitation, optical depth, chemistry, radiation, and ISM phase alter the response of a tracer, and the resulting differences can depend on both position and spatial scale. We propose scale-vector alignment, a scale-aware method based on Constrained Diffusion Decomposition (CDD). CDD decomposes an image into localized scale components; at each position, their amplitudes define a scale vector that describes how the measured intensity is distributed over spatial scale. We define the pixel-wise similarity \Spix(x,y)\Spix(x,y) as the normalized alignment of two local scale vectors. The normalization removes the overall amplitude, so \Spix\Spix compares relative scale composition rather than absolute flux. We also define the scale-wise similarity \Sscale(l)\Sscale(l) by comparing the two CDD component maps at each spatial scale. Spatial shifts are used to construct an empirical shifted reference distribution for \Spix\Spix. In OrionA, the tracer with the highest similarity to the dust-derived column-density map changes from 12^{12}CO to 13^{13}CO to C18^{18}O toward higher column density. In NGC6334I(N), the line--continuum similarity decreases locally around the brightest compact structures, where radiative-transfer effects can alter the observed line morphology. In NGC3627, CO is most similar to 21μμm emission, and \Sscale\Sscale reaches its maximum at an intermediate sub-kpc scale. The method measures where two tracers have similar multiscale structure and at which scales their spatial distributions agree. The implementation is publicly available at https://github.com/meng-ke/Scale-Vector-Alignment.
Sep 15, 2026cs.CV

G3AR: Graph-Guided Neural Visual Geometry for Scalable Multi-Sequence Aerial Registration

Full-context neural visual geometry is impractical for thousands of images, while sequence-based chunking poorly captures irregular non-local overlap in multi-sequence aerial collections. We present Graph-Guided Neural Visual Geometry for Aerial Registration (G3AR), a graph-guided framework for scalable dense neural geometry. Before local inference, G3AR builds a geometrically verified image-proximity graph that guides bounded overlapping chunks and induces a chunk graph whose maximum spanning tree defines alignment topology. Compatible backbones process chunks independently; shared-image predictions then estimate three-dimensional similarity (Sim(3)) transforms that register local cameras and geometry in a common frame. Across four real aerial scenes, G3AR improves pose error and runtime in matched VGGT- and Pi3-backed comparisons, while its DA3 variant achieves the lowest pose error among evaluated neural-geometry methods.
Sep 8, 2026cs.CV

RoMa-ΩΩ: What Feed-Forward 3D Models Know About Image Matching

Learned image matching has experienced significant progress in recent years, culminating in robust and accurate matchers such as RoMa, whose robustness is often attributed to its use of frozen DINO features. In a parallel development, feed-forward reconstruction models, such as VGGT, have been trained on ever-growing datasets to accurately regress dense 3D point maps and camera poses. The distinction between matchers and feed-forward reconstruction models has become increasingly blurred with the introduction of matching losses in models such as MASt3R and VGGT-ΩΩ. This raises a natural question: what do feed-forward 3D models know about image matching? In this work, we answer this question by analyzing three scenarios: (i) zero-shot matching of patch features, (ii) direct matching of 3D point predictions, and (iii) training a full matcher on top of the learned representations. We find that, despite performing poorly in zero-shot matching, especially in later layers, feed-forward reconstruction models provide strong representations for linear probing and full matching pipelines. We further show that, even without any training, their raw predictions alone enable competitive matching, albeit only under moderate viewpoint changes and modality gaps. Based on these insights, we retrain RoMa v2 by replacing its DINO backbone with VGGT-ΩΩ. Our resulting model, \ours, outperforms state-of-the-art matchers on a wide range of benchmarks, e.g. +8.1 mAA compared to RoMa v2 on WxBS.
Sep 2, 2026cs.CV

POI-Loc: A Fine-Grained POI Localization Benchmark and an Asymmetric Global-to-Local Matching Method

Point-of-interest (POI) localization matches user-provided storefront close-ups to the same shops in wide, geo-tagged vehicle-mounted street views. POIs may change while the surrounding scene stays similar, so scene-level recognition alone cannot establish POI identity. Differences in target scale and capture domains further challenge matching. We introduce POI-Loc, to our knowledge the first benchmark dedicated to this asymmetric, fine-grained POI localization task. Many visual place recognition methods represent each image with a single global vector, which tends to dilute fine-grained features of small storefronts amid background clutter. We propose GLAM (Global-to-Local Asymmetric Matching) to combine global and local evidence. In stage one, a single attention-pooled query probe is matched against compact reference region tokens via learnable soft top-k interaction, with the resulting local similarity fused with global similarity for retrieval. Stage two reuses query region tokens before attention pooling and stored reference tokens for mutual-nearest-neighbor re-ranking. GLAM surpasses both global and two-stage baselines on Recall@1/5/10 and mAP, with about 5×5\times smaller re-ranking features and 280×280\times lower per-pair matching cost than FoL. The benchmark and code will be released at https://github.com/roadhan/glam.
Aug 20, 2026cs.CV

RIPE++: Reinforced Keypoint Learning from Positive Pairs Only

Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure-from-motion, visual SLAM, augmented reality, and medical image registration. Learning robust local feature representations, however, typically requires accurate camera poses or depth supervision, which are often unavailable in real-world settings. Reinforcement learning (RL) has recently emerged as a promising alternative, requiring only the information if two images show the same scene or not. However, existing RL formulations such as RIPE rely on coarse binary rewards and carefully constructed negative training pairs, limiting training stability and descriptor discriminability. In this paper, we revisit RL-based keypoint learning and propose a reward that fully exploits the geometric consistency signal, deriving both reward and penalty from a single positive pair without contrasting against negatives. This richer signal provides sufficient supervisory contrast to learn discriminative detectors and descriptors from positive image pairs alone, enabling representation learning under extremely limited supervision. Furthermore, we show that the same RL objective can be extended to the matching stage by adapting LightGlue, raising AUC@5 on MegaDepth1500 from 56.58 to 59.65 and enabling weakly-supervised training of the full sparse matching pipeline from image pairs with partial visual overlap. We validate our approach on established benchmarks, demonstrating competitive results compared to fully-supervised methods. We further show that the method can be even trained on low texture medical video sequences, where camera poses are usually unavailable and standard SfM pipelines often fail. Code and data are available at https://github.com/fraunhoferhhi/RIPEpp .
Aug 13, 2026cs.CV

Online Learning of Correspondences between Images

We propose a novel method for iterative learning of point correspondences between image sequences. Points moving on surfaces in 3D space are projected into two images. Given a point in either view, the considered problem is to determine the corresponding location in the other view. The geometry and distortions of the projections are unknown as is the shape of the surface. Given several pairs of point-sets but no access to the 3D scene, correspondence mappings can be found by excessive global optimization or by the fundamental matrix if a perspective projective model is assumed. However, an iterative solution on sequences of point-set pairs with general imaging geometry is preferable. We derive such a method that optimizes the mapping based on Neyman's chi-square divergence between the densities representing the uncertainties of the estimated and the actual locations. The densities are represented as channel vectors computed with a basis function approach. The mapping between these vectors is updated with each new pair of images such that fast convergence and high accuracy are achieved. The resulting algorithm runs in real-time and is superior to state-of-the-art methods in terms of convergence and accuracy in a number of experiments.
Aug 12, 2026eess.IV

Alignment of Similarity-Transformed Images Based on Fourier--Mellin Transform Using Auxiliary Function Method

This paper proposes an algorithm for estimating the similarity transformation, namely translation, scale, and rotation, between two images with subpixel accuracy. Image registration is a fundamental technique for aligning images acquired under different viewpoints and imaging conditions, and a representative approach based on maximizing discrete cross-correlation is the Fourier--Mellin registration. However, the Fourier--Mellin approach often fails to achieve sufficient alignment accuracy when subpixel-level estimation is required. The proposed method integrates (i) scale-and-rotation estimation from the Fourier magnitude spectrum in a log-polar representation and (ii) maximization of phase-only correlation based on the auxiliary function method. This integration enables a two-stage estimation procedure: it first estimates scale and rotation without being affected by translation, and then estimates translation with subpixel precision in the spatial domain using the corrected image pair. A simulation experiment on image pairs subjected to random similarity transformations demonstrates that the proposed method reduces estimation errors in scale, rotation, and translation compared with Fourier--Mellin-based registration methods using discrete cross-correlation.
Aug 11, 2026cs.LG

Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the emergence of vision foundation models (VFMs). Despite these advances, existing studies remain highly diverse in their problem formulations, model architectures, training paradigms, and evaluation protocols, making it difficult to obtain a unified understanding of the field. In this survey, we present a unified review of cross-view feature matching. We first introduce a structured taxonomy covering feature extraction, single-type feature matcher, multi-type feature matcher, VFMs based methods, training strategy and robust estimation, providing a coherent framework for analysis and comparison. We further examine recent advances, distilling key design principles and highlighting the shift toward unified and generalizable correspondence models. We also provide a unified experimental benchmarking of representative state-of-the-art methods under consistent protocols, enabling fair and comprehensive performance comparisons. In addition, we discuss open challenges and future directions, including efficiency, robustness under extreme conditions, and cross-domain generalization. This survey aims to provide a comprehensive and structured reference for understanding the evolution, current landscape, and future development of cross-view feature matching in the era of vision foundation models.
Aug 11, 2026cs.CV

Multi-Level Evidence Aggregation for Robust Facial Phenotype Retrieval in Rare Genetic Disorder Prioritization

AI-assisted facial phenotyping supports rare genetic disorder prioritization by retrieving visually similar diagnosed cases from facial image reference databases such as the GestaltMatcher Database (GMDB). Existing GestaltMatcher-based retrieval frameworks compare each test image with individual gallery images in a facial phenotype embedding space. However, this pointwise formulation does not fully exploit available evidence, because patients may have multiple images and disorders may be represented by multiple diagnosed gallery patients. We propose an inference-time multi-level evidence aggregation framework that improves facial phenotype retrieval without modifying the underlying GestaltMatcher-Arc encoder. The framework combines embedding-level patient aggregation of multiple images from the same individual, patient-weighted disorder centroids, and hybrid individual-centroid scoring to integrate test-patient observations, disorder-level gallery evidence, and local nearest-neighbor evidence. We evaluated the approach on GMDB v1.1.4 across disorders represented during training (GMDB-Freq), unseen disorders (GMDB-Rare), and multi-image patient subsets, using a unified gallery containing both GMDB-Freq and GMDB-Rare disorders. Multi-level evidence aggregation improved mean per-disorder top-NN retrieval accuracy across all evaluation subsets. Top-1 accuracy increased from 38.52% to 48.82% on GMDB-Freq and from 19.38% to 23.79% on GMDB-Rare. On multi-image subsets, top-1 accuracy increased from 46.12% to 60.94% on GMDB-Multi-Freq and from 18.54% to 26.71% on GMDB-Multi-Rare. These findings show that inference-time aggregation can improve next-generation facial phenotype retrieval without retraining the encoder, supporting a shift from isolated single-image matching toward multi-level aggregation of patient and disorder evidence for rare-disorder prioritization.
Aug 10, 2026cs.CV

TeaMatch: Teachable Cross-Modal Representation Learning for 2D-3D Matching

Learning reliable correspondences between images and point clouds is fundamental for 2D-3D matching. Despite recent progress in detection-free methods, existing approaches primarily optimize matching within a single model and often struggle to maintain reliable correspondences under challenging conditions such as noisy inputs, low overlap, and ambiguous structures. In this work, we propose TeaMatch, a novel framework that introduces teachability as a criterion for cross-modal representation learning. We define teachability as the ability of a representation to be effectively recovered by weak learners under degraded inputs, reflecting its structural consistency and robustness. To this end, we construct a set of task-specific weak students that simulate common failure modes and train them to imitate the teacher on a training split while evaluating their recoverability on a disjoint meta split. The teacher is then optimized to improve the students' ability to recover reliable correspondences, guided by correspondence-level and geometry-aware constraints. Our framework can be seamlessly integrated into existing coarse-to-fine matching pipelines without additional inference cost. Extensive experiments demonstrate that TeaMatch improves matching robustness and achieves state-of-the-art performance on challenging 2D-3D matching benchmarks.
Aug 10, 2026cs.CV

XFeat Revisited: Reproducibility and Evaluation of a Lightweight Image Matcher

We present a reproducibility study of XFeat, a lightweight local feature extractor and matcher designed to identify corresponding points across images efficiently on resource-constrained hardware. We re-implement the architecture based on the paper and supplementary material, re-evaluate the authors' released checkpoint alongside our re-implementation, and conduct additional architectural ablations to examine design choices that were not fully justified in the original work. This distinction between re-evaluation and reproduction is important, as the paper, supplement, and public code differ in several implementation details, including the backbone layout, fusion block, and training losses. Empirically, our reproduced models closely match and, in some cases, outperform the re-evaluated original checkpoint on MegaDepth-1500 and ScanNet-1500, supporting the main claim that XFeat provides a strong accuracy-efficiency trade-off for standard image-matching benchmarks. Our ablations provide a more nuanced view of two architectural arguments from the original paper. In particular, the parallel keypoint branch is important for semi-dense matching, but its benefit is less pronounced than originally claimed, while the evidence for the specific placement of the single skip-connection remains inconclusive. Finally, we reproduce the original downstream evaluations and find close agreement for homography estimation, while Aachen visual localization remains below the reported results, even for the released checkpoint, suggesting sensitivity to underspecified evaluation details. We then extend the analysis to zero-shot out-of-distribution and cross-modal matching across retinal, thermal-visible, and multimodal remote-sensing imagery, where XFeat remains effective in some settings but degrades sharply under severe modality shifts.
Aug 9, 2026cs.CV

Semi-Dense Matching Uncertainty Is Not Just Local Confidence

Reliable semi-dense matching is essential for modern geometric vision systems. Designed under a coarse-to-fine paradigm, it achieves an optimal balance between performance and computational cost. However, existing methods often struggle to provide well-quantified uncertainties, where catastrophic coarse-assignment failures are ignored, leading to truncated error distributions and severely misjudged geometric estimations. In this paper, we propose a lightweight, post-hoc overall uncertainty estimation framework that introduces a two-component calibrated Laplace mixture model with only 9 learnable parameters. The objective is to explicitly capture both the sharp local refinement noise and the broader tail of coarse-assignment failures. We introduce the Coarse-success posterior Refit (CoRe) method, a geometric refitting module that utilizes the posterior probability of coarse-assignment success as soft correspondence weights. Extensive experiments show that our method consistently improves downstream geometric accuracy across various pretrained-only matchers and robust estimators with minimal computational overhead. Our code is available at https://github.com/khoavpt/Probabilistic-matching.
Aug 4, 2026cs.CV

LoRetta: A Foundation Model and Extensive Dataset for Global-Scale Remote Sensing Dense Image Matching

Dense image matching establishes pixel-wise correspondences and underpins broad applications in computer vision and photogrammetry. However, extending dense matching to global-scale remote sensing remains challenging because image pairs may differ in acquisition time, season, viewpoint, spatial resolution, and land-cover state. The resulting large geometric offsets, partial overlap, and intrinsically unmatchable regions make direct dense correspondence prediction unreliable and inefficient. We thus reformulate dense matching as localization-and-registration: first localizing the matchable overlap and affine geometry, then refining dense residuals within the aligned frame. Based on this formulation, we propose LoRetta, a foundation model coupling matchability-aware affine localization with guided dense registration. We also introduce LEVIR-GM, a global-scale multi-temporal optical matching benchmark with dataset-native matchability labels (103K aligned, 827K augmented pairs, six continents, five years, 0.5-1024 m resolution). We further establish a unified evaluation protocol for sparse, semi-dense, and dense matchers. On LEVIR-GM, LoRetta achieves an area under the curve (AUC) of 83.3%, outperforming the strongest baseline RoMa v2 by 1.6 points, with larger percentage of correct keypoints (PCK) gains of 6.5 and 8.2 points at 1 and 2 pixels, while reducing inference latency by 47.8%. Astronaut-to-satellite and unmanned aerial vehicle (UAV)-to-satellite geolocalization experiments further demonstrate its transferability as a reusable geometric aligner.
Aug 4, 2026cs.CV

SGFormer: Structure-Guided Transformer for Robust Local Feature Matching

Local feature matching is a fundamental component of photogrammetry, enabling accurate image correspondence critical for tasks such as 3D reconstruction, stereo mapping, and visual localization. While recent detector-free matching methods, like LoFTR, have advanced the field, the global features obtained by leveraging the global-range modeling capacity of the unconstrained attention mechanism compromise the model's attention to the salient structures in certain scenarios. This limitation leads to a phenomenon we define as attention divergence, wherein a portion of high-confidence matches are distributed outside the valid matching region (overlapping region), especially in scenes with large viewpoint variations. This occurs because similar features in irrelevant regions may receive equal weighting and consideration within the standard Transformer, limiting matching reliability in challenging photogrammetric environments. To address this issue in feature matching, we propose SGFormer (Structure-Guided Transformer), a novel structure-aware matching network that adaptively updates attention on features near salient structure in overlapping regions. SGFormer employs a semi-dense coarse-to-fine pipeline and incorporates the proposed Triple-Structure-Attention (TSA) module into the backbone net for extracting distinctive features. The TSA module utilizes shallow local features from early network layers to enhance the representation around salient structure, guiding subsequent transformer stages to intensify the model's focus on regions with salient structure across the global scope. SGFormer, thereby reinforcing attention to visually consistent areas while mitigating the influence of non-overlapping regions. Extensive experiments show that SGFormer significantly mitigates attention divergence and improves matching accuracy.
Aug 4, 2026cs.CV

Double Down on Defense: Strengthening Deep Perceptual Hashes against Evasion Attacks without Retraining

Near-duplicate image matching is crucial for trust and safety, provenance verification, copyright enforcement, and large-scale visual search. Modern platforms increasingly rely on deep perceptual hashes, which map visually similar images to nearby representations despite common image transformations. However, adversarial perturbations can cause near-duplicates to evade matching. We present DualShield, a plug-in defense that improves the robustness of existing deep perceptual hashes without retraining or modifying their underlying models. DualShield combines matching-time randomized smoothing, which aggregates decisions over perturbed reference-query pairs, with publication-time hardening, which adds an optimized imperceptible perturbation to each reference image before publication. Together, these mechanisms provide certified and empirical robustness. DualShield achieves a certified ℓ2\ell_2 radius of approximately 0.3, guaranteeing that query perturbations within this radius cannot evade matching. We further evaluate it against adaptive white-box, black-box, and image-transformation attacks. Across eight deep perceptual hashes and three datasets, DualShield substantially reduces attack success rates while preserving low collision rates. These results show that deep perceptual hashes can be strengthened without costly retraining by improving the matching procedure and hardening reference images before publication.
Aug 3, 2026cs.CV

Calibrated Similarity and Graph Clustering for Open-Set Animal Re-Identification

AnimalCLEF26 addresses discovery-oriented animal re-identification, where systems must both attach query images to known individuals and discover unseen individuals by clustering them correctly. We present a similarity-to-clustering pipeline for this setting across Eurasian lynx, fire salamander, loggerhead sea turtle, and Texas horned lizard images. The method first isolates the target specimen using segmentation and then applies lightweight species-specific preprocessing for lynx, sea turtle, and salamander images to enhance identity-relevant visual cues, while Texas horned lizard images are used after segmentation only. Pairwise similarities are then estimated with WildFusion by calibrating and combining a MiewID global descriptor with two local matching branches, ALIKED + LightGlue and DISK + LightGlue. The resulting query-query similarities are refined and converted into identity clusters using graph-based clustering, while query-database similarities are used to attach confident samples to known identities. We evaluate training-free and fine-tuned MiewID variants, including Dynamic ArcFace and SphereFace2-Focal adaptations, and combine them in the final ensemble. Our selected ensemble substantially improves on the WildFusion baseline, achieving the best public ARI of 0.72124 and a private ARI of 0.70393, while a simpler preprocessing-before-calibration variant achieves the best private ARI of 0.71087. These results indicate that calibrated global-local fusion with species-aware preprocessing choices is effective for open-set wildlife re-identification under challenging field conditions and visual variation. The implementation code is available on GitHub.
Jul 29, 2026cs.CV

Robust RPC Bundle Adjustment for Multi-Date Satellite Imagery with Season-Invariant Correspondences

Accurate refinement of Rational Polynomial Camera (RPC) models is essential for high-quality satellite image geolocation. In ground control point (GCP)-free multi-view pipelines, this refinement is commonly performed through bundle adjustment from automatically extracted image correspondences. However, conventional RPC bundle adjustment pipelines rely on handcrafted feature matching, which becomes unreliable in multi-date collections affected by seasonal, illumination, and land-cover changes. We propose an appearance-aware RPC refinement pipeline that combines learned local feature matching for season-invariant correspondences with global image descriptors for selecting visually compatible image pairs. This reduces redundant and error-prone matching while preserving the connectivity of the matching graph. Experiments on seasonally diverse WorldView-3 images show that our pipeline improves GCP-free relative RPC refinement over open-source baselines, achieving lower geometric consistency errors while substantially reducing matching time on collections with 39-42 views. By making RPC refinement more robust to diachronic appearance variation, our approach enables more effective use of multi-date satellite imagery.
Jul 28, 2026cs.CV

HOME: Robust Hough-space Matching Method for Structured and Textureless Videos

Visual front-ends for robotic localization typically rely on point-based features such as Oriented FAST and Rotated BRIEF (ORB), which frequently fail in structured environments dominated by strong linear structures or textureless surfaces. While line-based Simultaneous Localization and Mapping (SLAM) systems mitigate this by utilizing line segments, conventional line extraction and description algorithms are computationally prohibitive for real-time edge robotics. To address this fundamental bottleneck, we propose HOME (Hough-space One-dimensional Matching of Extrema), an ultra-lightweight, training-free feature matching framework. HOME transforms images into Hough space, mapping global linear structures to stable local extrema, which serve as keypoints, thereby reformulating complex line matching into highly efficient one-dimensional point matching. The proposed 1D radial descriptor mathematically guarantees rotational and translational invariance without the overhead of explicit orientation estimation. As a proof of concept to validate the matching accuracy and efficiency of HOME, this paper focuses on homography estimation. Extensive evaluations demonstrate that HOME achieves robust registration in challenging scenarios where point-based methods fail, operating at a much faster speed than existing line-based methods. Extending this robust matching engine to full 3D pose estimation remains a highly promising future direction.
Jul 21, 2026cs.CV

SynGallery: A Synthetic Gallery of Real Paintings for Instance-Level Artwork Recognition

Instance-level artwork recognition requires matching a handheld visitor photograph to a specific work in a large museum collection. This is challenging because painting datasets typically provide clean catalog images for training, while test queries are captured under oblique viewpoints, gallery lighting, reflections, frames, and other scene-level variations. We present SynGallery, a synthetic gallery dataset for artwork retrieval that addresses this gap without collecting additional real photographs. Starting from catalog images of real paintings, we place each artwork into a procedurally generated 3D gallery scene and render it from multiple viewpoints under varied geometric and appearance conditions, while preserving the exact identity of the original work. The resulting dataset contains 24,490 rendered views of 4,898 paintings from the Met benchmark. We show that these synthetic views provide a stronger training signal than the corresponding studio photographs. At the same number of training data points, training only on SynGallery improves art painting recognition from 67.18 to 73.47 GAP−^-. When added to the full Met training set, SynGallery improves the published benchmark protocol from 35.97 to 38.48 GAP. Ablation experiments show that the gain comes primarily from geometric viewpoint variation rather than photographic realism: blur, sensor noise, and image compression consistently reduce performance.
Jul 17, 2026cs.CV

DS@GT ARC at AnimalCLEF 2026: Species-Aware Graph Construction for Multi-Species Animal Re-Identification

Automated individual animal re-identification is essential for large-scale biodiversity monitoring; however, field imagery complicates separating identity cues from nuisance variation in pose, illumination, background, resolution, and species-specific morphology. The DS@GT ARC submission to AnimalCLEF 2026 introduces a multi-species image-clustering system for re-identifying Eurasian lynx, fire salamanders, loggerhead sea turtles, and Texas horned lizards. Instead of relying on a single descriptor or nearest-neighbor retrieval, this approach formulates re-identification as species-aware graph construction over candidate image pairs. The pipeline integrates tailored preprocessing, global candidate retrieval, LightGlue-based local verification with multiple keypoint families, LightGBM pair scoring, conservative edge admission, and Leiden community detection. This design directly addresses a primary failure mode of clustering-based re-identification: high-scoring false pairs that act as bridge edges and merge distinct individuals through transitive closure. Across species, ablation studies demonstrate that local feature support, foreground-aware preprocessing, and species-specific backbone selection enhance pair evidence, while graph operating points determine the trade-off between fragmentation and over-merging. The selected submission achieved a public ARI of 0.733 and a private ARI of 0.674, ranking fifth among 230 teams. These results indicate that robust wildlife re-identification requires not only strong visual representations but also calibrated integration of global similarity, local identity markings, neighborhood context, and graph-level constraints. The code can be found at https://github.com/dsgt-arc/animalclef-2026.
Jul 16, 2026cs.CV

Hough-SIFT: Robust Image Registration for Linear Structures via Hough Space

Image registration is essential in applications such as electronic image stabilization. Scale-Invariant Feature Transform (SIFT), a widely used local keypoint detector and descriptor, typically provides accurate registration; however, it often fails in scenes with strong linear structures (e.g., shutters), where local features become ambiguous. We propose Hough-SIFT, a robust registration method that performs SIFT descriptor matching in Hough space. In this domain, linear structures form distinctive peaks that restore descriptor discriminability. Experiments demonstrate that Hough-SIFT is robust in linear scenes where SIFT frequently fails, while maintaining accuracy comparable to SIFT in normal scenes.
Jul 7, 2026cs.CV

RFHNet: Relational and Frequency-Aware Hashing Network for Large-Scale Fine-Grained Food Image Retrieval

Fine-grained food image retrieval is a key task in computational gastronomy, with applications in food traceability, dietary monitoring, and smart catering systems. Although hashing-based retrieval is attractive for large-scale search due to its storage efficiency and fast Hamming-distance computation, existing methods often perform poorly in fine-grained food scenarios, where subtle local semantics and frequency-sensitive visual cues are essential. To address this challenge, we propose RFHNet, a cascaded hierarchical hashing network that captures both global structure and fine-grained local details through multi-level representations. RFHNet includes three components: (1) Fine-grained Relation Modeling (FRM) to capture subtle visual differences among similar food components; (2) Multi-Frequency Modulated Fusion (MFMF) to extract informative multi-frequency features; and (3) Hierarchical Semantic Synergy (HSS) to adaptively integrate multi-level representations and generate discriminative hash codes. Experiments on six food-specific benchmarks show that RFHNet consistently outperforms state-of-the-art hashing methods, with mAP gains of 4.44% to 17.20% at 12 bits. These results validate the effectiveness of RFHNet for large-scale visual food retrieval and smart catering applications. The source code will be released upon publication.
Jul 6, 2026cs.CV

Hybrid Deep Learning for Traceability and Classification of Industrial Slate Tiles

Applying deep learning to instance-aware reidentification of slate tiles and extraction site classification can improve production efficiency and quality control in the slate tile industry. These tasks are particularly important for handling natural materials where visual variability can make manual inspection costly and error-prone. We present a lightweight, hybrid deep learning approach that combines image matching and classification within a single framework. The system integrates a feature-matching branch based on XFeat with a MobileNetV3- based classification branch. The XFeat branch, combined with a LightGlue matching head, improves instance matching performance by +15.4% AUC. For classification, features from both backbones are shared and fused, resulting in a +10.9% accuracy improvement over a standard MobileNetV3 model. Our approach is evaluated on a newly created industrial dataset consisting of 2,610 slate tile images from six extraction sites. The results demonstrate the effectiveness of the proposed approach for object re-identification and classification in an industrial setting.