Robust person re-identification often combines complementary cues such as face, gait, and body shape. While adaptive fusion typically targets query quality, model strength also varies across identities. We introduce identity-conditioned score fusion, a framework that tailors weights to each gallery identity without training. By contrasting intra-identity consistency against cross-identity impostors, it extracts identity-specific profiles that couple with query-conditioned adaptation via a parameter-free rule. This widens the separation between true and false matches while preserving score calibration. Evaluations on three clothes-changing person re-identification benchmarks show that our method consistently outperforms statistical, rank-based, and learned baselines, achieving up to an 8.8% absolute reduction in the false non-identification rate and demonstrating the value of identity-conditioned fusion in open-set person re-identification.
Figures & tables
Figure 1: Query-conditioned vs. Identity-conditioned score fusion. Existing multi-modal gait recognition often applies a global, shared weight to all identities (left). In contrast, our proposed Identity-conditioned approach adaptively tailors weights for each identity profile (right), allowing the system to leverage the most reliable modality on an identity-specific basis. Sliders conceptually illustrate the weight assignment.
Figure 2: Overview of our framework. Top: Our method leverages gallery intra- and inter-identity statistics to derive tailored model reliability weights wki for each identity i without auxiliary training. Bottom: When an incoming query is compared against the gallery, its similarity scores with candidate identity i across modalities are combined using that candidate’s specific profile, prioritizing reliable modalities for each individual for final open-set verification.
Dataset
Type
# Identities
# Queries
# Gallery exemplars
Train / Test
Total
Mean / identity
Range / identity
CCVID
Video
75 / 151
834
1,074
7.1
3–12
MEVID
Video
104 / 54
316
1,438
26.6
1–101
LTCC
Image
77 / 75
493
7,050
94.0
4–589
Table 1: Statistics of person re-identification benchmarks. Query and gallery counts refer to video sequences for CCVID ( Gu et al., 2022b ) and MEVID ( Davila et al., 2023 ) , and images for LTCC ( Qian et al., 2020 ) . Identity counts are reported as train/test. The mean and range of gallery exemplar counts are computed across gallery identities in the test split.
Table 2: Performance on CCVID and MEVID. Comb. specifies the model pool shared by each block ( ⧫ AdaFace, ♠ CAL, ▲ BigGait on CCVID and AGRL on MEVID, ♣ InsightFace/ArcFace, and ⊛ LoRA-adapted CLIP-ReID); the top block reports individual-model results and the two below it the fusion methods. Best and second-best fusion results are marked within each pool; highlighted rows are ours. Values are in %; FNIR is mean ± standard deviation over ten enrollment trials. R-1: Rank-1 accuracy, TAR: TAR@1%FAR, FNIR: FNIR@1%FPIR.
Figure 3: Identity profile permutation on CCVID. Expert pool: AdaFace, CAL, BigGait, and InsightFace. Gray histograms show performance distributions over 1,000 random identity-to-profile permutations while keeping models, gallery, and weight vectors identical. The genuine assignment (solid orange line) strictly outperforms all permutations across all three metrics, as well as equal weighting (dashed line), confirming that gains stem from identity-specific weights rather than arbitrary non-uniform weighting.
Table 6
Figure 5: Robustness to gallery enrichment. Queries are added into the gallery under their predicted identity i^ when the calibrated match score satisfies Sq,i^>η . Solid curves and shaded regions denote the mean and spread across 9 arrival sequences, while dashed lines mark the static-gallery baselines. We report: (a) TAR@ 1% FAR and (b) number of admitted queries across enrollment thresholds η . Ours (vanilla) consistently outperforms both equal weighting and QME across all thresholds, remaining robust regardless of enrichment aggressiveness.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Separation of match and non-match score distributions. Dataset: CCVID; model pool: AdaFace, BigGait, and CAL. Dashed lines indicate distribution means and the 1% FAR threshold. Our identity-conditioned fusion markedly enlarges the margin between genuine matches and impostors by driving match scores rightward, effectively suppressing distribution overlap under strict verification thresholds.
Combination Rule
Rank-1
mAP
TAR@1%FAR
FNIR@1%FPIR ( ↓ )
Query only
93.77
91.96
82.40
10.38
Identity only
93.53
93.10
85.36
11.23
Product (Default)
94.60
93.99
85.60
9.96
Weighted sum ( λ=0.25 )
94.00
92.94
84.45
10.27
Weighted sum ( λ=0.50 )
93.89
93.38
85.94
10.51
Weighted sum ( λ=0.75 )
93.64
93.31
86.08
10.72
Appendix
Table A1: Ablation on weight combination rules for joint fusion. Dataset: CCVID; Models pool: AdaFace, BigGait, CAL, and zero-shot CLIP-ReID. We compare different strategies for composing query-conditioned weights wkq and identity-specific weights wki , keeping each branch’s weights fixed. λ denotes the interpolation weight for the identity branch. Bold marks the best result among all combination strategies.
CCVID
MEVID
Approach
R-1 ↑
mAP ↑
TAR ↑
FNIR ↓
R-1 ↑
mAP ↑
TAR ↑
FNIR ↓
Equal weights (no identity score)
92.03
90.62
80.72
15.96
62.13
36.62
40.37
58.92
d′ (Default)
94.21
92.62
84.36
10.60
64.03
38.10
41.75
57.57
Mean genuine similarity
93.65
91.76
83.09
9.67
61.50
36.10
40.12
59.97
Genuine − impostor mean
93.94
91.95
83.34
9.74
61.71
36.44
40.23
59.82
AUC
92.18
90.76
82.66
16.46
63.40
37.89
41.57
57.84
Appendix
Table A2: Ablation on per-identity separability metrics. Comparison of different gallery-derived separation scores and learned per-identity weighting baselines ( ‡ ), evaluated across four CCVID and three MEVID expert pools. FNIR and TAR are measured at 1% FPIR and 1% FAR, respectively. Bold indicates the best result per column.
Figure A2: Robustness to gallery enrichment with a margin-based admission rule. Queries are enrolled when the gap between the top two candidate scores satisfies Sq,i^−maxi=i^Sq,i>η . Across all subplots, curves and shaded areas indicate the mean and spread over 9 arrival orders, and dashed lines show static-gallery baselines: (a) TAR@ 1% FAR, (b) pseudo-label precision of admitted queries, and (c) total admitted queries ( η=0 admits all queries). Ours (vanilla) consistently maintains its performance advantage over equal weighting and QME at every threshold.
Combination
Method
Rank-1
mAP
TAR@1%FAR
FNIR@1%FPIR ( ↓ )
⧫
AdaFace ( Kim et al., 2022 )
94.0
87.9
75.7
13.0 ± 3.5
♠
CAL ( Gu et al., 2022a )
81.4
74.7
66.3
52.8 ± 13.3
▲
BigGait ( Ye et al., 2024 )
76.7
61.0
49.7
71.1 ± 6.1
♣
InsightFace/ArcFace ( Deng et al., 2019 )
94.1
88.2
78.9
18.1 ± 3.1
⚫
CLIP-ReID ( Li et al., 2023a )
64.1
63.0
67.9
63.1 ± 2.5
⊛
LoRA-adapted CLIP-ReID ( Hu et al., 2022 )
85.4
84.3
80.3
42.8 ± 7.5
Appendix
Table A4: Extended multi-model performance on CCVID. Combination specifies the model pool shared by each block ( ⧫ AdaFace, ♠ CAL, ▲ BigGait, ♣ InsightFace/ArcFace ( Deng et al., 2019 ) , and ⚫ zero-shot and ⊛ LoRA-adapted CLIP-ReID ( Li et al., 2023a ; Hu et al., 2022 ) ). Best and second-best results are marked within each model pool.
Combination
Method
Rank-1
mAP
TAR@1%FAR
FNIR@1%FPIR ( ↓ )
⧫
AdaFace ( Kim et al., 2022 )
25.0
8.1
5.4
98.8 ± 1.2
♠
CAL ( Gu et al., 2022a )
52.5
27.1
34.7
67.8 ± 7.3
▲
AGRL ( Wu et al., 2020 )
51.9
25.5
30.7
69.4 ± 8.9
⚫
CLIP-ReID ( Li et al., 2023a )
63.0
33.0
33.6
68.3 ± 3.3
⊛
LoRA-adapted CLIP-ReID ( Hu et al., 2022 )
69.9
45.1
47.7
49.7 ± 4.3
Min-Fusion ( Jain et al., 2005 )
46.8
21.2
28.0
70.4 ± 8.0
Appendix
Table A5: Extended multi-model evaluation on MEVID. Combination specifies the model pool shared by each block ( ⧫ AdaFace, ♠ CAL, ▲ AGRL, and ⚫ zero-shot and ⊛ LoRA-adapted CLIP-ReID ( Li et al., 2023a ; Hu et al., 2022 ) ). Best and second-best results are marked within each block.
Combination
Method
Rank-1
mAP
TAR@1%FAR
FNIR@1%FPIR ( ↓ )
⧫
AdaFace ( Kim et al., 2022 )
18.5
5.9
2.4
99.8 ± 0.2
♠
CAL ( Gu et al., 2022a )
74.4
40.6
36.7
59.7 ± 7.3
■
AIM ( Yang et al., 2023 )
74.8
40.9
37.0
66.2 ± 7.5
⚫
CLIP-ReID ( Li et al., 2023a )
70.7
33.7
26.5
47.8 ± 13.0
⊛
LoRA-adapted CLIP-ReID ( Hu et al., 2022 )
76.7
40.7
31.9
42.0 ± 10.3
Min-Fusion ( Jain et al., 2005 )
37.1
12.9
12.4
81.9 ± 6.0
Appendix
Table A6: Extended multi-model performance on LTCC. Combination specifies the model pool shared by each block ( ⧫ AdaFace, ♠ CAL, ■ AIM, and ⚫ zero-shot and ⊛ LoRA-adapted CLIP-ReID ( Li et al., 2023a ; Hu et al., 2022 ) ). Best and second-best results are marked within each model pool.
Learning identity-discriminative representations with multi-scene generality has become a critical objective in person re-identification (ReID). However, mainstream perception-driven paradigms tend to identify fitting from massive annotated data rather than identity-causal cues understanding, which presents a fragile representation against multiple disruptions. In this work, ReID-R is proposed as a novel reasoning-driven paradigm that achieves explicit identity understanding and reasoning by incorporating chain-of-thought into the ReID pipeline. Specifically, ReID-R consists of a two-stage contribution: (i) Discriminative reasoning warm-up, where a model is trained in a CoT label-free manner to acquire identity-aware feature understanding; and (ii) Efficient reinforcement learning, which proposes a non-trivial sampling to construct scene-generalizable data. On this basis, ReID-R leverages high-quality reward signals to guide the model toward focusing on ID-related cues, achieving accurate reasoning and correct responses. Extensive experiments on multiple ReID benchmarks demonstrate that ReID-R achieves competitive identity discrimination as superior methods using only 14.3K non-trivial data (20.9% of the existing data scale). Furthermore, benefit from inherent reasoning, ReID-R can provide high-quality interpretation for results.
Domain Generalizable (DG) person re-identification (Re-ID) has attracted growing research interest due to its potential for deployment in unseen real-world scenarios. Most existing approaches address DG Re-ID by focusing on training domain-generalizable encoders but ignore the possible refinements in inference stage. In contrast, this work explores an alternative direction which improves inference re-ranking to enhance DG Re-ID. Conventional re-ranking methods typically rely on neighborhood-based distances to refine the initial ranking list, inherently depending on features produced by the Re-ID encoder. However, they deteriorate on target domains since the encoder lacks sufficient generalizability to produce reliable feature distances on unseen scenarios. Inspired by the remarkable generalization capabilities of recent Multimodal Large Language Models (MLLMs), we propose an MLLM-empowered distance metric to improve re-ranking in DG Re-ID. Specifically, we first adapt an MLLM to Re-ID data through supervised fine-tuning, which incorporates a domain-agnostic prompt and a query-candidate hard mining scheme. Then, the adapted MLLM is employed to compute a μ-distance during inference, which is robust to domain gap and significantly enhances subsequent re-ranking performance. Our approach is model-agnostic and can be seamlessly integrated into previous re-ranking frameworks. Extensive experiments demonstrate that our approach consistently yields substantial performance improvements across multiple DG Re-ID benchmarks. The code of this work will be released at https://github.com/RikoLi/MUSE soon.
Jiachen Li, Xiaojin Gong
College of Information Science and Electronic Engineering Zhejiang University
CLIP-based person re-identification (ReID) methods aggregate spatial features into a single global \texttt{[CLS]} token optimized for image-text alignment rather than spatial selectivity, making representations fragile under occlusion and cross-camera variation. We propose SAGA-ReID, which reconstructs identity representations by aligning intermediate patch tokens with anchor vectors parameterized in CLIP's text embedding space -- emphasizing spatially stable evidence while suppressing corrupted or absent regions, without requiring textual descriptions of individual images. Controlled experiments isolate the aggregation mechanism under two qualitatively distinct conditions -- synthetic masking, where identity signal is absent, and realistic human distractors, where an overlapping person introduces semantically confusing signal -- with SAGA's advantage over global pooling growing substantially as occlusion increases across both conditions. Benchmark evaluations confirm consistent gains over CLIP-ReID across standard and occluded settings, with the largest improvements where global pooling is most unreliable: up to +10.6 Rank-1 on occluded benchmarks. SAGA's aggregation outperforms dedicated sequential patch aggregation on a stronger backbone, confirming that structured reconstruction addresses a bottleneck that backbone quality and architectural complexity alone cannot resolve. Code available at https://github.com/ipl-uw/Structured-Anchor-Guided-Aggregation-for-ReID.
Aotian Zheng, Winston Sun, Bahaa Alattar +2
Department of Electrical and Computer Engineering, University of Washington, Seattle, WA 98195 · Applied Physics Laboratory, University of Washington, Seattle, WA 98195