Organizations: Faculty of Information Science and Engineering, Ocean University of China, Qingdao, China · Sanya Oceanographic Institution, Ocean University of China, Sanya, China
Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arises in animal ReID through markings, contours, and scars, in person ReID through subtle clothing and accessory cues, and in vehicle ReID through localized appearance details. Although visual foundation models encode such information in dense tokens, a single holistic descriptor can obscure discriminative local signals. We propose FM-ReID, an end-to-end framework that formulates local representation learning as selective competitive token routing. Its Competitive Fine-grained Mining module uses multiple mining queries and a residual query to compete for dense DINOv3 tokens. Above-prior selection retains tokens preferentially allocated to each mining query, while the residual slot receives tokens excluded from the retrieval descriptors. The resulting multi-query descriptors are jointly trained with a holistic representation for retrieval, without fixed spatial partitions or equal-area constraints. FM-ReID achieves strong results on animal, person, and vehicle ReID benchmarks, supporting competitive token routing as an effective way to augment holistic foundation-model representations.
Figures & tables
Figure 1: (a) Object ReID often requires localized identity cues to complement holistic representations when different instances have similar global appearances. (b) FM-ReID selectively mines such cues from dense foundation-model tokens through competitive query routing, without fixed spatial partitions or equal-area constraints.
Figure 2: Overview of FM-ReID. DINOv3 provides a holistic CLS token and dense patch tokens. CFM selectively routes patch-token evidence through competing mining and residual queries, then retains above-prior assignments for the mining descriptors. During training, the holistic feature, holistic–fine-grained pairs, and their joint fusion receive ReID supervision. At inference, normalized holistic and fine-grained features form the retrieval embedding, while the residual output is discarded.
Figure 3: Competitive Fine-grained Mining (CFM). Mining queries and a residual query compete for each patch token through a softmax across queries. Above-prior mining assignments are retained and renormalized over tokens to form query-specific descriptors; their supports may overlap and vary in size. Query refinement produces the fine-grained features, whereas the residual output is excluded from the retrieval embedding.
Model
Input
mTop-1
mTop-5
BAKS
ConvNeXt-Base
2242
81.1
90.3
78.5
EfficientNet-B3
3002
77.8
88.5
75.1
ViT-Base
2242
78.3
88.7
75.8
Swin-Base
2242
81.5
90.4
79.1
TransReID
2562
80.52
89.46
78.22
CLIP-ReID
2562
76.49
87.61
74.21
Table 1: Comparison on (a) WildlifeReID-10k under the official closed-set protocol and (b) vehicle ReID benchmarks. The first four baselines in (a) are from Adam et al. (2025) ; TransReID and CLIP-ReID are from our runs. Bold indicates the best result in each metric.
Method
Reference
Backbone
MSMT17
Market-1501
DukeMTMC
Occ-Duke
mAP
R1
mAP
R1
mAP
R1
mAP
R1
PCB Sun et al. (2018)
ECCV18
CNN
40.4
68.2
81.6
93.8
69.2
83.3
-
-
OSNet Zhou et al. (2019)
ICCV19
52.9
78.7
84.9
94.8
73.5
88.6
-
-
Fast-ReID He et al. (2023)
MM23
59.9
83.3
-
-
78.9
89.6
-
-
TransReID He et al. (2021)
ICCV21
ViT
67.4
85.3
88.9
95.2
82.0
90.7
59.2
66.4
DCAL Zhu et al. (2022a)
CVPR22
64.0
83.1
87.5
94.7
80.1
89.0
-
-
Table 2: Comparison with selected published methods on MSMT17, Market-1501, DukeMTMC-reID, and Occ-Duke. Bold indicates the best result among the listed methods.
Source
IDs
Images
Self-collected
36
1,205
FS-48 subset
48
22,692
AAUZebraFish
5
281
WhaleSharkID
519
7,305
MFT25
116
14,149
Total
724
45,632
Table 3: FM-FISH: (a) dataset statistics and (b) comparison results. MFT25 identity labels denote annotated tracks. Bold indicates the best result in each metric.
Backbone
Holistic
CFM
WildlifeReID-10k
MSMT17
CLS
GAP
Mining
Residual
mTop-1
mTop-5
BAKS
mAP
R1
DINOv3-B/16
✓
83.29
91.44
81.46
76.1
89.6
✓
✓
84.28
91.56
82.41
77.4
90.5
✓
✓
✓
84.54
92.01
82.90
79.6
91.4
✓
✓
✓
✓
85.03
92.07
83.33
79.7
91.4
Table 4: Component ablation on WildlifeReID-10k and MSMT17. Mining and Residual denote the mining queries and auxiliary residual query. Within each benchmark, all variants use DINOv3-B/16, the same training recipe, and seed 1234; CFM variants retain the default fusion and diversity objectives. Bold indicates the best result in each metric.
Figure 4: Sensitivity to the number of mining queries K : (a) WildlifeReID-10k, with BAKS/mTop-1 on the left/right axes; (b) MSMT17, with mAP/Rank-1 on the left/right axes. Dotted lines mark K=5 ; the residual query is retained and excluded from K . All runs use epoch-120 checkpoints and seed 1234. Panel (a) retains the full-model result at K=5 , while (b) uses independent sweep runs for every K .
Method
Parameters (M)
GFLOPs
Peak memory (MiB)
TransReID
92.92
40.77
371.5
CLIP-ReID
86.14
22.78
346.1
Baseline
85.66
23.40
341.5
Baseline + GAP
86.84
23.40
346.0
FM-ReID
(+9.66%) 95.23
(+1.41%) 23.73
(+9.28%) 378.1
Table 5: Inference cost at 256×128 , FP32, and batch size 1 on an RTX 5080. Baseline uses DINOv3 with CLS pooling; parentheses indicate increases over Baseline + GAP.
Figure 5: Qualitative visualization on two panda identities. From left to right: input images, holistic responses, and five miner attention maps. Holistic responses use feature–patch cosine similarity, while miner maps show token aggregation weights. Each map is independently min–max normalized for display; colors indicate relative responses within each map.
As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancements have been made in object ReID. However, most existing methods suffer from generalization due to the limited size and diversity of ReID datasets. Meanwhile, current models tend to focus on extracting semantic patterns rather than learning identity-aware feature distributions. To address these issues, we propose a novel feature learning framework named \textbf{DiffReID} for object ReID. It leverages a discriminative diffusion model to gradually learn identity-aware distributions and generate identity-invariant features. More specifically, with the Contrastive Language-Image Pre-training (CLIP) model, we first obtain identity-aware text features by prompt tuning. Then, we propose a Vision-guided Noise Generator (VNG) to initialize probabilistic noises and gradually corrupt identity-aware text features. Afterwards, we take visual features as conditions and propose a Light Weight Denoiser (LWD) to denoise the corrupted text features step-by-step for identity-aware distribution learning. To obtain discriminative features, we further generate identity-invariant guided features from randomly sampling visual-guided noises. Finally, we propose a Mutual Enhancement Constraint (MEC) to facilitate mutual learning between visual features and guided features to enhance the representation robustness and discrimination. Extensive experiments on five object ReID benchmarks demonstrate that our method shows better results than most state-of-the-art methods. The source code is available at https://github.com/AWangYQ/DiffReID.
Yingquan Wang, Pingping Zhang, Dong Wang +1
School of Information and Communication Engineering, Dalian University of Technology · School of Future Technology, Dalian University of Technology
Learning identity-discriminative representations with multi-scene generality has become a critical objective in person re-identification (ReID). However, mainstream perception-driven paradigms tend to identify fitting from massive annotated data rather than identity-causal cues understanding, which presents a fragile representation against multiple disruptions. In this work, ReID-R is proposed as a novel reasoning-driven paradigm that achieves explicit identity understanding and reasoning by incorporating chain-of-thought into the ReID pipeline. Specifically, ReID-R consists of a two-stage contribution: (i) Discriminative reasoning warm-up, where a model is trained in a CoT label-free manner to acquire identity-aware feature understanding; and (ii) Efficient reinforcement learning, which proposes a non-trivial sampling to construct scene-generalizable data. On this basis, ReID-R leverages high-quality reward signals to guide the model toward focusing on ID-related cues, achieving accurate reasoning and correct responses. Extensive experiments on multiple ReID benchmarks demonstrate that ReID-R achieves competitive identity discrimination as superior methods using only 14.3K non-trivial data (20.9% of the existing data scale). Furthermore, benefit from inherent reasoning, ReID-R can provide high-quality interpretation for results.
Any-Time Person Re-identification (AT-ReID) necessitates the robust retrieval of target individuals under arbitrary conditions, encompassing both modality shifts (daytime and nighttime) and extensive clothing-change scenarios, ranging from short-term to long-term intervals. However, existing methods are highly relying on pure visual features, which are prone to change due to environmental and time factors, resulting in significantly performance deterioration under scenarios involving illumination caused modality shifts or cloth-change. In this paper, we propose Semantic-driven Token Filtering and Expert Routing (STFER), a novel framework that leverages the ability of Large Vision-Language Models (LVLMs) to generate identity consistency text, which provides identity-discriminative features that are robust to both clothing variations and cross-modality shifts between RGB and IR. Specifically, we employ instructions to guide the LVLM in generating identity-intrinsic semantic text that captures biometric constants for the semantic model driven. The text token is further used for Semantic-driven Visual Token Filtering (SVTF), which enhances informative visual regions and suppresses redundant background noise. Meanwhile, the text token is also used for Semantic-driven Expert Routing (SER), which integrates the semantic text into expert routing, resulting in more robust multi-scenario gating. Extensive experiments on the Any-Time ReID dataset (AT-USTC) demonstrate that our model achieves state-of-the-art results. Moreover, the model trained on AT-USTC was evaluated across 5 widely-used ReID benchmarks demonstrating superior generalization capabilities with highly competitive results. Our code will be available soon.
Jiaxuan Li, Xin Wen, Zhihang Li
Tsinghua University · Independent Researcher · University of Chinese Academy of Sciences