Organizations: Faculty of Information Science and Engineering, Ocean University of China, Qingdao, China · Sanya Oceanographic Institution, Ocean University of China, Sanya, China
Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arises in animal ReID through markings, contours, and scars, in person ReID through subtle clothing and accessory cues, and in vehicle ReID through localized appearance details. Although visual foundation models encode such information in dense tokens, a single holistic descriptor can obscure discriminative local signals. We propose FM-ReID, an end-to-end framework that formulates local representation learning as selective competitive token routing. Its Competitive Fine-grained Mining module uses multiple mining queries and a residual query to compete for dense DINOv3 tokens. Above-prior selection retains tokens preferentially allocated to each mining query, while the residual slot receives tokens excluded from the retrieval descriptors. The resulting multi-query descriptors are jointly trained with a holistic representation for retrieval, without fixed spatial partitions or equal-area constraints. FM-ReID achieves strong results on animal, person, and vehicle ReID benchmarks, supporting competitive token routing as an effective way to augment holistic foundation-model representations.
Figures & tables
Figure 1: (a) Object ReID often requires localized identity cues to complement holistic representations when different instances have similar global appearances. (b) FM-ReID selectively mines such cues from dense foundation-model tokens through competitive query routing, without fixed spatial partitions or equal-area constraints.
Figure 2: Overview of FM-ReID. DINOv3 provides a holistic CLS token and dense patch tokens. CFM selectively routes patch-token evidence through competing mining and residual queries, then retains above-prior assignments for the mining descriptors. During training, the holistic feature, holistic–fine-grained pairs, and their joint fusion receive ReID supervision. At inference, normalized holistic and fine-grained features form the retrieval embedding, while the residual output is discarded.
Figure 3: Competitive Fine-grained Mining (CFM). Mining queries and a residual query compete for each patch token through a softmax across queries. Above-prior mining assignments are retained and renormalized over tokens to form query-specific descriptors; their supports may overlap and vary in size. Query refinement produces the fine-grained features, whereas the residual output is excluded from the retrieval embedding.
Model
Input
mTop-1
mTop-5
BAKS
ConvNeXt-Base
2242
81.1
90.3
78.5
EfficientNet-B3
3002
77.8
88.5
75.1
ViT-Base
2242
78.3
88.7
75.8
Swin-Base
2242
81.5
90.4
79.1
TransReID
2562
80.52
89.46
78.22
CLIP-ReID
2562
76.49
87.61
74.21
Table 1: Comparison on (a) WildlifeReID-10k under the official closed-set protocol and (b) vehicle ReID benchmarks. The first four baselines in (a) are from Adam et al. (2025) ; TransReID and CLIP-ReID are from our runs. Bold indicates the best result in each metric.
Method
Reference
Backbone
MSMT17
Market-1501
DukeMTMC
Occ-Duke
mAP
R1
mAP
R1
mAP
R1
mAP
R1
PCB Sun et al. (2018)
ECCV18
CNN
40.4
68.2
81.6
93.8
69.2
83.3
-
-
OSNet Zhou et al. (2019)
ICCV19
52.9
78.7
84.9
94.8
73.5
88.6
-
-
Fast-ReID He et al. (2023)
MM23
59.9
83.3
-
-
78.9
89.6
-
-
TransReID He et al. (2021)
ICCV21
ViT
67.4
85.3
88.9
95.2
82.0
90.7
59.2
66.4
DCAL Zhu et al. (2022a)
CVPR22
64.0
83.1
87.5
94.7
80.1
89.0
-
-
Table 2: Comparison with selected published methods on MSMT17, Market-1501, DukeMTMC-reID, and Occ-Duke. Bold indicates the best result among the listed methods.
Source
IDs
Images
Self-collected
36
1,205
FS-48 subset
48
22,692
AAUZebraFish
5
281
WhaleSharkID
519
7,305
MFT25
116
14,149
Total
724
45,632
Table 3: FM-FISH: (a) dataset statistics and (b) comparison results. MFT25 identity labels denote annotated tracks. Bold indicates the best result in each metric.
Backbone
Holistic
CFM
WildlifeReID-10k
MSMT17
CLS
GAP
Mining
Residual
mTop-1
mTop-5
BAKS
mAP
R1
DINOv3-B/16
✓
83.29
91.44
81.46
76.1
89.6
✓
✓
84.28
91.56
82.41
77.4
90.5
✓
✓
✓
84.54
92.01
82.90
79.6
91.4
✓
✓
✓
✓
85.03
92.07
83.33
79.7
91.4
Table 4: Component ablation on WildlifeReID-10k and MSMT17. Mining and Residual denote the mining queries and auxiliary residual query. Within each benchmark, all variants use DINOv3-B/16, the same training recipe, and seed 1234; CFM variants retain the default fusion and diversity objectives. Bold indicates the best result in each metric.
Figure 4: Sensitivity to the number of mining queries K : (a) WildlifeReID-10k, with BAKS/mTop-1 on the left/right axes; (b) MSMT17, with mAP/Rank-1 on the left/right axes. Dotted lines mark K=5 ; the residual query is retained and excluded from K . All runs use epoch-120 checkpoints and seed 1234. Panel (a) retains the full-model result at K=5 , while (b) uses independent sweep runs for every K .
Method
Parameters (M)
GFLOPs
Peak memory (MiB)
TransReID
92.92
40.77
371.5
CLIP-ReID
86.14
22.78
346.1
Baseline
85.66
23.40
341.5
Baseline + GAP
86.84
23.40
346.0
FM-ReID
(+9.66%) 95.23
(+1.41%) 23.73
(+9.28%) 378.1
Table 5: Inference cost at 256×128 , FP32, and batch size 1 on an RTX 5080. Baseline uses DINOv3 with CLS pooling; parentheses indicate increases over Baseline + GAP.
Figure 5: Qualitative visualization on two panda identities. From left to right: input images, holistic responses, and five miner attention maps. Holistic responses use feature–patch cosine similarity, while miner maps show token aggregation weights. Each map is independently min–max normalized for display; colors indicate relative responses within each map.