BagDINO: Multi-View Baggage Re-Identification with DINOv3
Authors: Vita Santa Barletta, Danilo Caivano, Rebecca Margiotta, Massimiliano Morga, Davide Pio Posa
Organizations: University of Bari Aldo moro Department of Computer Science Bari, Italia · University of Bari Aldo Moro, SER&Practices Department of Computer Science Bari, Italia · SER&Practices Spin-off of the University of Bari Aldo Moro Bari, Italia
Mishandled checked baggage remains a recurrent issue in airport operations, and current recovery workflows still largely rely on tag-based tracking, which does not directly support visual identification when tag evidence is missing or unavailable. This paper investigates baggage re-identification as an instance-level retrieval problem in a multi-camera setting, leveraging DINOv3 foundation-model representations to match a query image against a gallery of registered baggage images. A Torchreid-style BNNeck re-identification head is placed on top of a DINOv3 backbone, and parameter-efficient adaptation is performed via LoRA. Experiments are conducted on the MVB benchmark using a progressive study that compares a fully frozen backbone against LoRA and fine-tuning strategies. Results indicate that parameter-efficient adaptation of foundation-model features provides an effective and stable approach for multi-view baggage re-identification under limited training data.
Figures & tables
Fig. 1: Overview of the proposed baggage re-identification pipeline. A DINOv3 ViT backbone provides a global image representation (from the [CLS] token) and is adapted via LoRA modules injected into selected transformer projections (e.g., qkv , proj , fc1 , fc2 ) while keeping the pre-trained weights frozen. The resulting embedding is processed by a Torchreid-style BNNeck head and optimized with a composite objective that combines identity classification (cross-entropy with label smoothing) and batch-hard triplet loss. At inference time, the linear classifier is discarded and retrieval is performed using the normalized embeddings (e.g., cosine similarity) to rank gallery items for a given query.
Fig. 2: Inference-stage retrieval. Query and gallery images are embedded by the backbone+BNNeck, and gallery instances are ranked by similarity in the embedding space to produce the final match list.
Fig. 3: Example MVB query-gallery pairs illustrating the heterogeneity of baggage items and the variability induced by different acquisition conditions.
Fig. 4: High-similarity failure cases on MVB. Three examples are reported, each shown as (i) the query, (ii) the corresponding true match in the gallery (positive), and (iii) the top-ranked incorrect retrieval returned with high confidence (similarity ≳78% ). These errors typically occur when different baggage identities exhibit near-duplicate visual cues (e.g., similar shape, material, and color patterns), yielding very close neighbors in the embedding space.
Method
mAP
Rank-1
BagDINO (Frozen)
0.57
0.60
BagDINO + LoRA
0.88
0.87
BagDINO FT (25%)
0.83
0.83
BagDINO FT (100%)
0.83
0.83
TABLE I: Results on the MVB benchmark in terms of mAP and Rank-1 summarizing the evaluated DINOv3-based configurations under different adaptation strategies.
Architecture
mAP
Rank-1
Huangbin et al. [ 14 ]
0.83
0.85
Yang et al. [ 9 ]
0.86
0.88
Liu et al. [ 10 ]
0.87
0.86
Zhao et al. [ 15 ]
0.91
0.87
Mazzeo et al. [ 8 ]
0.63
0.69
Zhiwei et al. [ 11 ]
0.90
0.89
TABLE II: Results on the MVB benchmark in terms of mAP and Rank-1 comparing BagDINO best approach with the representative methods from literature.
Method
mAP
Rank-1
BagDINO + LoRA
93.63
96.32
Yang et al. [ 9 ]
88.60
95.40
Mazzeo et al. [ 8 ]
80.92
85.79
TABLE III: Cross-domain evaluation on Market-1501 to assess the transferability of the proposed approach beyond the baggage domain.
Sun Yat-sen University, China · Pazhou Lab (HuangPu), Guangdong, China · Guangdong Province Key Laboratory of Information Security Technology, Guangzhou, China +1