WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification
Authors: Turhan Can Kargin, Piotr Kubaty, Ekaterina Rostovskaya, Izabela Wierzbowska, Bartosz Zieliński, Marcin Przewięźlikowski
Organizations: Faculty of Mathematics and Computer Science, Jagiellonian University, Poland · Doctoral School of Exact and Natural Sciences, Jagiellonian University, Poland · Institute of Environmental Sciences, Faculty of Biology, Jagiellonian University, Kraków, Poland · Jagiellonian Center for Artificial Intelligence, Kraków, Poland · NASK National Research Institute, Warsaw, Poland
Individual animal re-identification from camera-trap imagery is an instance retrieval problem central to non-invasive wildlife monitoring: a query image must retrieve the correct individual from a reference set of known animals. This requires computer vision models to recognize distinctive local patterns in fur, skin, or other visual markings. Current approaches either learn global embeddings as a classification problem, requiring many labeled images per individual while largely ignoring local evidence, or apply off-the-shelf, domain-agnostic image matchers. Although such matchers are pretrained on large and diverse image collections, adapting them to wildlife imagery is challenging because available datasets are small and lack correspondence-level annotations. We study weakly supervised adaptation of a pretrained keypoint matcher using only identity labels, without keypoint-level or geometric correspondence ground truth. We mine informative image pairs with the pretrained matcher, derive weak positive and negative supervision from identity agreement, and contrastively fine-tune the matching network to strengthen correspondences for same-identity pairs and suppress them for different identities. Across open-source wildlife re-identification datasets, our approach improves accuracy over off-the-shelf matchers and a state-of-the-art local--global fusion method. Under an open-world protocol with held-out individuals, it learns a transferable correspondence prior rather than memorizing training identities. To our knowledge, this is the first study of matcher-level, identity-supervised adaptation for animal re-identification. Our method enables data-efficient specialization of image matching models to wildlife domains using identity annotations already available in typical monitoring datasets.
Figures & tables
Figure 1: Wildlife re-identification as instance retrieval. A camera-trap query is compared with every known individual, and a local feature matcher ranks them by the summed confidence of point pairs that both images agree on, giving a short list for an expert. The database holds hundreds of individuals, of which six are shown. WildMatch performs weakly supervised adaptation of the matcher using only the identity labels of the database (Fig. 2 ), without keypoint or correspondence annotation. Images are synthetic CzechLynx renders [ 20 ] , with illustrative matches and scores.
Figure 2: Weakly supervised matcher adaptation with WildMatch. Before training, the pretrained matcher selects top-scoring positives and hard negatives. A triplet loss then fine-tunes the matcher to score positive pairs above hard negatives.
Table 1: Dataset statistics. Database (DB) images form the retrieval gallery and are the only images used for fine-tuning, and query (Q) images are evaluated against them. Med. and Max give the median and maximum number of database images per individual. The small plot shows the image counts of all individuals, sorted from most to least photographed, with the top 10% in dark blue, and Top 10% gives their share of all database images. Single is the share of individuals with one image.
Figure 3: Examples of challenging images in the benchmark datasets. Raw frames that are overexposed (a), show too little of the animal to see its markings (b), contain no visible animal (c), are corrupted (d), or are blurred (e). Such images remain in the database and query sets of all methods.
Figure 4: Retrieval accuracy on eight datasets. Top-5 accuracy as a function of the candidate budget k (log scale). LoMa + WildMatch (ours) is compared with the default LoMa matcher, WildFusion, and cosine retrieval with MegaDescriptor-L, which also provides the candidates for all matching methods, and with DINOv3-L (flat lines).
LoMa
RDD-LightGlue
Fine-tuned part
Top-5
Bal.
Δ
GPU-h
Top-5
Bal.
Δ
GPU-h
none (default)
55.2
31.3
–
–
56.0
33.4
–
–
descriptor branch
56.1
28.6
− 2.6
77
49.8
25.7
− 7.7
89
matching module
58.3
34.7
+3.4
5.1
56.8
34.4
+0.9
5.0
both
59.9
34.9
+3.6
84
59.0
35.7
+2.2
103
Table 2: Descriptor or matching module. For each matcher, the descriptor branch, the matching module, or both are fine-tuned with the same identity supervision and training data (CzechLynx, closed split, k=250 ). Bal. is balanced top-1 accuracy, Δ its change from the default checkpoint, and GPU-h the training cost in GPU-hours of training steps. The matching-module row (ours) is shaded.
Figure 5: WildMatch and classification by training cost (CzechLynx, closed split, k=250 ). Top-5 (left) and balanced top-1 accuracy (right) of LoMa + WildMatch checkpoints and of the MegaDescriptor-L classifier with the backbone frozen, partially fine-tuned (partial FT), or fully fine-tuned (full FT), as a function of GPU-hours on the same hardware. The default LoMa matcher and cosine retrieval need no training and are drawn as flat lines.
k=10
k=50
k=100
k=160
Method
Top-5
Bal.
Top-5
Bal.
Top-5
Bal.
Top-5
Bal.
MegaDescriptor-L cosine
19.6
12.0
19.6
12.0
19.6
12.0
19.6
12.0
DINOv3-L cosine
20.2
11.1
20.2
11.1
20.2
11.1
20.2
11.1
WildFusion
21.7
17.3
27.7
21.1
30.6
22.5
32.6
23.1
LoMa (default)
22.7
18.9
30.9
23.0
35.5
25.8
40.4
26.7
LoMa + WildMatch
25.0
19.7
36.6
27.3
40.9
30.3
46.1
31.8
Table 3: Unseen-identity protocol on CzechLynx (Sec. 4 ). Top-5 and balanced top-1 accuracy (%) at candidate budgets k . Cosine retrieval does not depend on k . Fine-tuned matchers (ours) are shaded; best value per column in bold.
Figure 6: Generality to a second matcher. Top-5 accuracy of RDD + WildMatch (ours) and the default RDD-LightGlue matcher as a function of the candidate budget k (log scale) on three datasets, with WildFusion and cosine retrieval for reference. The advantage of fine-tuning is largest on Nyala and smaller on Salamander and CzechLynx.
Figure 7: Qualitative results of top-1 retrievals with LoMa + WildMatch . For each dataset, a query (left) and its top-1 gallery image (right) of the same individual. Of the 350 to 420 matches found for each pair, lines show the 10 most confident. Matching uses background-removed inputs, and the matches are drawn on the original photos.
AnimalCLEF26 addresses discovery-oriented animal re-identification, where systems must both attach query images to known individuals and discover unseen individuals by clustering them correctly. We present a similarity-to-clustering pipeline for this setting across Eurasian lynx, fire salamander, loggerhead sea turtle, and Texas horned lizard images. The method first isolates the target specimen using segmentation and then applies lightweight species-specific preprocessing for lynx, sea turtle, and salamander images to enhance identity-relevant visual cues, while Texas horned lizard images are used after segmentation only. Pairwise similarities are then estimated with WildFusion by calibrating and combining a MiewID global descriptor with two local matching branches, ALIKED + LightGlue and DISK + LightGlue. The resulting query-query similarities are refined and converted into identity clusters using graph-based clustering, while query-database similarities are used to attach confident samples to known identities. We evaluate training-free and fine-tuned MiewID variants, including Dynamic ArcFace and SphereFace2-Focal adaptations, and combine them in the final ensemble. Our selected ensemble substantially improves on the WildFusion baseline, achieving the best public ARI of 0.72124 and a private ARI of 0.70393, while a simpler preprocessing-before-calibration variant achieves the best private ARI of 0.71087. These results indicate that calibrated global-local fusion with species-aware preprocessing choices is effective for open-set wildlife re-identification under challenging field conditions and visual variation. The implementation code is available on GitHub.
Made In Alexandria Artificial Intelligence Team, Alexandria, Egypt · Faculty of Computers and Data Science, Alexandria University, Alexandria, Egypt · Faculty of Computer Science and Engineering, Alamein International University, New Alamein City, 51718, Egypt +3
Wildlife re-identification aims to recognise individual animals by matching query images to a database of previously identified individuals, based on their fine-scale unique morphological characteristics. Current state-of-the-art models for multispecies re- identification are based on deep metric learning representing individual identities by fea- ture vectors in an embedding space, the similarity of which forms the basis for a fast automated identity retrieval. Yet very often, the discriminative information of individual wild animals gets significantly reduced due to the presence of several degradation factors in images, leading to reduced retrieval performance and limiting the downstream eco- logical studies. Here, starting by showing that the extent of this performance reduction greatly varies depending on the animal species (18 wild animal datasets), we introduce an augmented training framework for deep feature extractors, where we apply artificial but diverse degradations in images in the training set. We show that applying this augmented training only to a subset of individuals, leads to an overall increased re-identification performance, under the same type of degradations, even for individuals not seen during training. The introduction of diverse degradations during training leads to a gain of up to 8.5% Rank-1 accuracy to a dataset of real-world degraded animal images, selected using human re-ID expert annotations provided here for the first time. Our work is the first to systematically study image degradation in wildlife re-identification, while introducing all the necessary benchmarks, publicly available code and data, enabling further research on this topic.
Thanos Polychronou, Lukáš Adam, Viktor Penchev +1
School of Mathematical Sciences, Queen Mary University of London, United Kingdom · University of West Bohemia in Pilsen, FEE, RICE, Pilsen, Czechia
Animal re-identification (ReID) in camera-trap surveys remains challenging due to low image quality, strong variation in illumination and viewpoint, and highly imbalanced numbers of observations per individual. As a result, current ReID performance is often insufficient for fully automated use, and practical workflows typically depend on expert review of algorithmically proposed candidate matches. Moreover, most existing approaches focus almost exclusively on visual cues and overlook auxiliary information routinely available in field studies, such as image timestamps and camera-trap locations. We introduce Spotted, a location-informed, human-in-the-loop animal ReID framework that integrates visual similarity with spatio-temporal feasibility priors derived from camera locations, thereby reducing the amount of required expert review. Our method (i) computes an image-model-agnostic feasibility score based on the minimum travel speed required for two detections to correspond to the same individual, (ii) uses these feasibility cues as pseudo-supervision to train a lightweight head on top of a frozen visual foundation model, and (iii) fuses adapted visual similarity with spatio-temporal feasibility to obtain a robust pairwise matching score. We additionally integrate an active pair sampling strategy to accelerate annotation by initially prioritizing uncertain predictions. We evaluate Spotted on three challenging camera-trap ReID datasets comprised of spotted hyenas and leopards, which we release as part of this work. Our model improves average top-5 identification accuracy by 9pp, 2pp and 9pp over the best baseline on our LeopardID102, SpottedHyenaID109 and SpottedHyenaID415 datasets, respectively. Further, we show that our human-in-the-loop strategy reduces the number of queried comparisons by up to 69pp while achieving equivalent positive matches.
Halil Sina Kelebek, Julia Hindel, Kobus Hoffman +9
Oxford Robotics Institute, University of Oxford, Oxford, UK · Department of Computer Science, University of Freiburg, Germany · Bubye Valley Conservancy, Zimbabwe +3