BagDINO: Multi-View Baggage Re-Identification with DINOv3
Authors: Vita Santa Barletta, Danilo Caivano, Rebecca Margiotta, Massimiliano Morga, Davide Pio Posa
Organizations: University of Bari Aldo moro Department of Computer Science Bari, Italia · University of Bari Aldo Moro, SER&Practices Department of Computer Science Bari, Italia · SER&Practices Spin-off of the University of Bari Aldo Moro Bari, Italia
Mishandled checked baggage remains a recurrent issue in airport operations, and current recovery workflows still largely rely on tag-based tracking, which does not directly support visual identification when tag evidence is missing or unavailable. This paper investigates baggage re-identification as an instance-level retrieval problem in a multi-camera setting, leveraging DINOv3 foundation-model representations to match a query image against a gallery of registered baggage images. A Torchreid-style BNNeck re-identification head is placed on top of a DINOv3 backbone, and parameter-efficient adaptation is performed via LoRA. Experiments are conducted on the MVB benchmark using a progressive study that compares a fully frozen backbone against LoRA and fine-tuning strategies. Results indicate that parameter-efficient adaptation of foundation-model features provides an effective and stable approach for multi-view baggage re-identification under limited training data.
Figures & tables
Fig. 1: Overview of the proposed baggage re-identification pipeline. A DINOv3 ViT backbone provides a global image representation (from the [CLS] token) and is adapted via LoRA modules injected into selected transformer projections (e.g., qkv , proj , fc1 , fc2 ) while keeping the pre-trained weights frozen. The resulting embedding is processed by a Torchreid-style BNNeck head and optimized with a composite objective that combines identity classification (cross-entropy with label smoothing) and batch-hard triplet loss. At inference time, the linear classifier is discarded and retrieval is performed using the normalized embeddings (e.g., cosine similarity) to rank gallery items for a given query.
Fig. 2: Inference-stage retrieval. Query and gallery images are embedded by the backbone+BNNeck, and gallery instances are ranked by similarity in the embedding space to produce the final match list.
Fig. 3: Example MVB query-gallery pairs illustrating the heterogeneity of baggage items and the variability induced by different acquisition conditions.
Fig. 4: High-similarity failure cases on MVB. Three examples are reported, each shown as (i) the query, (ii) the corresponding true match in the gallery (positive), and (iii) the top-ranked incorrect retrieval returned with high confidence (similarity ≳78% ). These errors typically occur when different baggage identities exhibit near-duplicate visual cues (e.g., similar shape, material, and color patterns), yielding very close neighbors in the embedding space.
Method
mAP
Rank-1
BagDINO (Frozen)
0.57
0.60
BagDINO + LoRA
0.88
0.87
BagDINO FT (25%)
0.83
0.83
BagDINO FT (100%)
0.83
0.83
TABLE I: Results on the MVB benchmark in terms of mAP and Rank-1 summarizing the evaluated DINOv3-based configurations under different adaptation strategies.
Architecture
mAP
Rank-1
Huangbin et al. [ 14 ]
0.83
0.85
Yang et al. [ 9 ]
0.86
0.88
Liu et al. [ 10 ]
0.87
0.86
Zhao et al. [ 15 ]
0.91
0.87
Mazzeo et al. [ 8 ]
0.63
0.69
Zhiwei et al. [ 11 ]
0.90
0.89
TABLE II: Results on the MVB benchmark in terms of mAP and Rank-1 comparing BagDINO best approach with the representative methods from literature.
Method
mAP
Rank-1
BagDINO + LoRA
93.63
96.32
Yang et al. [ 9 ]
88.60
95.40
Mazzeo et al. [ 8 ]
80.92
85.79
TABLE III: Cross-domain evaluation on Market-1501 to assess the transferability of the proposed approach beyond the baggage domain.
Aerial-Ground Person Re-Identification (AGPReID) remains highly challenging due to drastic viewpoint variations between drones and fixed cameras. Existing methods typically follow a view-invariant paradigm, aligning shared features across views to achieve robustness. However, view-invariant inherently enforces part-level alignment, which ignores view-specific cues and discriminative identity information. To this end, this work proposes ViSA (View-aware Semantic Alignment), a view-aware framework that achieves cross-view semantic consistency containing an Expert-driven Token Generation Module (ETGM) and a Dual-branch Local Fusion Module (DLFM). Technically, the former constructs a set of view-aware experts to generate adaptive semantic queries that perceive viewpoint-specific patterns, while the latter leverages graph reasoning to extract and align local regions responsive to different experts. Extensive experiments on three AGPReID benchmarks including AG-ReID.v2, CARGO and LAGPeR demonstrate that ViSA consistently achieves superior performance, with a notable 10.06% mAP improvement on the challenging CARGO cross-view protocol. The code is available at https://github.com/Cat-Zero/ViSA.
Quan Zhang, Zeqiang Cai, Peiming Zhao +4
Sun Yat-sen University, China · Pazhou Lab (HuangPu), Guangdong, China · Guangdong Province Key Laboratory of Information Security Technology, Guangzhou, China +1
As a fundamental image processing task, object Re-Identification (ReID) aims to retrieve objects across non-overlapping cameras. Recently, with the development of deep learning, significant advancements have been made in object ReID. However, most existing methods suffer from generalization due to the limited size and diversity of ReID datasets. Meanwhile, current models tend to focus on extracting semantic patterns rather than learning identity-aware feature distributions. To address these issues, we propose a novel feature learning framework named \textbf{DiffReID} for object ReID. It leverages a discriminative diffusion model to gradually learn identity-aware distributions and generate identity-invariant features. More specifically, with the Contrastive Language-Image Pre-training (CLIP) model, we first obtain identity-aware text features by prompt tuning. Then, we propose a Vision-guided Noise Generator (VNG) to initialize probabilistic noises and gradually corrupt identity-aware text features. Afterwards, we take visual features as conditions and propose a Light Weight Denoiser (LWD) to denoise the corrupted text features step-by-step for identity-aware distribution learning. To obtain discriminative features, we further generate identity-invariant guided features from randomly sampling visual-guided noises. Finally, we propose a Mutual Enhancement Constraint (MEC) to facilitate mutual learning between visual features and guided features to enhance the representation robustness and discrimination. Extensive experiments on five object ReID benchmarks demonstrate that our method shows better results than most state-of-the-art methods. The source code is available at https://github.com/AWangYQ/DiffReID.
Yingquan Wang, Pingping Zhang, Dong Wang +1
School of Information and Communication Engineering, Dalian University of Technology · School of Future Technology, Dalian University of Technology
Aerial-ground person re-identification is a challenging task due to cross-platform viewpoint variations, which cause severe occlusion and geometric deformation. Existing methods attempt to learn view-invariant representations exclusively within the 2D image space, where drastic viewpoint variations cause the learned features to remain coupled with viewpoint bias. To address this, we propose VR3D, a View-Robust 3D Representation Learning framework that maps images into a unified 3D coordinate space to achieve view-independent feature interaction. Specifically, we introduce View-Robust 3D Representation Interaction, which leverages 3D priors extracted from single 2D observations to lift 2D appearance features into a canonical 3D space. VR3I employs 3D Geometry-Semantic Attention to establish interactions between 2D patches and 3D voxels from corresponding body parts based on their 3D spatial locations, effectively grounding 2D semantics within a 3D framework. In addition, as the reliability of these representations varies across samples due to viewpoint changes and 3D reconstruction errors, we introduce Reliability-Aware Fusion, which estimates sample-specific reliability and adaptively aggregates the multi-source representations. Extensive experiments on three benchmark datasets (CARGO, AG-ReID.v1, and AG-ReID.v2) demonstrate that VR3D outperforms recent methods. For example, it achieves a 5.63% improvement in Rank-1 on CARGO. Our code will be released.
Chao Ji, Shiyu Xuan, Zechao Li
School of Computer Science and Engineering, Nanjing University of Science and Technology