Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
Authors: Ian de Holanda Cavalcanti Bezerra, Vivek Trivedy, Lucas Pascotti Valem, Longin Jan Latecki
Organizations: Institute of Mathematics and Computer Science (ICMC), University of S˜ao Paulo (USP), S˜ao Carlos, Brazil · Department of Computer and Information Science, Temple University, Philadelphia, PA, USA
Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers. However, trying to squeeze all the semantic information of an image into a single descriptor can hurt downstream retrieval performance, especially for fine-grained retrieval tasks. In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens. We leverage the DINOv2-reg model, which includes register tokens that emergently learn object and part-based representations. For each "cue" token ([CLS] and each register token), we find a "buddy" image patch token and extract an N x N patch region to produce a set of localized ROI tokens. Our approach automatically captures important regions of interest without any external bounding boxes or saliency modules, purely by matching semantic tokens with their spatial representation regions. Furthermore, we incorporate these tokens into a multi-vector retrieval framework inspired by ColBERT, enabling fine-grained matching via a per-token alignment mechanism while avoiding the large storage cost of keeping all patch embeddings. Through extensive experiments, we find that (1) register tokens encode useful fine-grained details that can complement the [CLS] token; (2) automatically pooled ROI tokens further improve fine-grained discrimination; and (3) multi-vector retrieval with a small set of tokens improves over a DINOv2-reg single-vector baseline while remaining tractable for large-scale search. The code is available at https://github.com/IdhcbIan/Augmenting_CLS_with_ROI_tokens.
Figures & tables
Fig. 1: Distribution of most similar tokens of the database images for any given [CLS] token from queries of images in COCO’s dog class [ 15 ] . Interestingly, a [CLS] token from a query image is often most similar to one of the four register tokens of other images (64%), not their [CLS] token (36%).
Fig. 2: For each cue token (the [CLS] and register tokens), the location of its corresponding buddy token is shown by the black max box. They focus on different parts of the semantic input: dog’s head, dog’s paw, cat’s ear, and ball.
Fig. 3: The data flow of our training setup. First, the [CLS] and register tokens are extracted with DINOv2-reg. Second, for each of these cue tokens, the buddy token is identified as the most similar image patch token. Third, the ROI token is constructed by mean pooling the buddy token and the image patch tokens in the region around it. In our experiments, we use only 10 tokens per image. Finally, the standard triplet training is performed with ColBERT-style similarity between image pairs.
Method
Dim.
Architecture
CUB-200
In-Shop
Cars-196
Stanford Online Products
1
2
4
8
1
10
20
30
1
2
4
8
1
2
4
8
NSoftmax [ 28 ]
512
ResNet50
61.3
73.9
83.5
90
86.6
97.5
98.4
98.8
84.2
90.4
94.4
96.9
78.2
90.6
96.2
-
ProxyNCA++ [ 27 ]
512
ResNet50
69.0
79.8
87.3
92.7
90.4
98.1
98.8
99.0
86.5
92.5
95.7
97.7
80.7
92.0
96.7
98.9
A-BIER [ 29 ]
512
GoogleNet
57.5
68.7
78.3
86.2
93.1
95.1
96.9
97.5
82.0
89.0
93.2
96.1
74.2
86.9
94.0
97.8
ABE [ 30 ]
512
GoogleNet
60.6
71.5
79.8
87.4
87.3
96.7
97.9
98.2
85.2
90.5
94.0
96.1
76.3
88.4
94.8
98.2
SM [ 31 ]
512
GoogleNet
56.0
68.3
78.2
86.3
90.7
97.8
98.5
98.8
83.4
89.9
93.9
96.5
75.3
87.5
93.7
97.4
TABLE I: Recall@k metrics on CUB-200, In-Shop, Cars-196, and Stanford Online Products. The table compares reported methods with different backbones, embedding dimensions, and training objectives; our DINOv2-reg ablations isolate the contribution of the proposed token selection and multi-vector scoring.
R-MAC [ 44 ]
GeM [ 45 ]
GeM+AP [ 46 ]
SSLeb [ 47 ]
43.0%
69.0%
33.5%
92.2%
SSLss [ 47 ]
SSLeb+Bd [ 47 ]
SSLss+Bd [ 47 ]
Ours
92.0%
92.5%
92.3%
93.5%
TABLE II: Comparison of retrieval performance (mAP) on the INSTRE dataset against state-of-the-art methods.
Method
CUB (R@1)
Cars-196 (R@1)
DINOv2-reg ( [CLS] only)
86.6
89.8
DINOv2-reg ( [CLS] +Registers)
86.7
90.3
DINOv2-reg ( [CLS] +Registers+ROI)
87.4
90.5
TABLE III: Effect of using register and ROI tokens (bottom 2 experiments using multi-vector training).
ROI Size
CUB (R@1)
Cars-196 (R@1)
Single Patch
87.8
89.4
3×3
88.1
90.1
5×5
88.1
90.5
7×7
88.2
90.0
9×9
87.8
90.1
TABLE IV: Ablation on the region size for ROI tokens. We report Recall@1 on CUB and Cars-196 with single patch vs. N×N mean pooling for N=3, 5, 7, 9. Our default setting is N=3.
Method
# Tokens
T. Dim
Memory
Recall@1
Single-Vector
1
384
1.5 GB
86.6
[CLS] +Registers
5
1,920
7.7 GB
86.7
Ours: [CLS] +Reg+ROI
10
3,840
15 GB
87.4
All Patch+Reg+ [CLS]
201
77,184
309 GB
87.3
TABLE V: Theoretical index size comparisons for 1M images using 384-dimensional float32 embeddings (4 bytes/dim). Recall@1 is measured on CUB-200 and does not represent million-scale retrieval performance. “T. Dim” and “Memory” denote dimensions per image and raw storage without overhead, respectively.
CUB-200
In-Shop
R@1
R@2
R@4
R@1
R@10
R@20
RRT [ 17 ]
68.7
85.0
95.6
88.3
97.9
98.6
LOCORE-tiny [ 48 ]
71.4
86.8
96.4
89.1
97.9
98.2
LOCORE-small [ 48 ]
74.6
89.1
97.3
89.4
97.7
97.7
LOCORE-base [ 48 ]
78.3
91.9
98.2
87.9
97.9
98.7
Ours
87.4
92.7
95.2
92.9
98.4
98.9
TABLE VI: Comparison of our approach with local feature-based methods on metric learning benchmarks in terms of Recall@k.