Region-Aware CLS Token Augmentation for Fine-Grained Image Retrieval
Authors: Ian de Holanda Cavalcanti Bezerra, Vivek Trivedy, Lucas Pascotti Valem, Longin Jan Latecki
Organizations: Institute of Mathematics and Computer Science (ICMC), University of S˜ao Paulo (USP), S˜ao Carlos, Brazil · Department of Computer and Information Science, Temple University, Philadelphia, PA, USA
Image retrieval methods often rely on a single global semantic descriptor extracted from an image, e.g., the [CLS] token in vision transformers. However, trying to squeeze all the semantic information of an image into a single descriptor can hurt downstream retrieval performance, especially for fine-grained retrieval tasks. In this work, we augment the semantic tokens in the newer visual transformers, the global [CLS] token and the four register tokens, with a carefully selected collection of spatial tokens, aiming to capture the spatial region representation that characterizes the contents captured in each of the semantic tokens. We leverage the DINOv2-reg model, which includes register tokens that emergently learn object and part-based representations. For each "cue" token ([CLS] and each register token), we find a "buddy" image patch token and extract an N x N patch region to produce a set of localized ROI tokens. Our approach automatically captures important regions of interest without any external bounding boxes or saliency modules, purely by matching semantic tokens with their spatial representation regions. Furthermore, we incorporate these tokens into a multi-vector retrieval framework inspired by ColBERT, enabling fine-grained matching via a per-token alignment mechanism while avoiding the large storage cost of keeping all patch embeddings. Through extensive experiments, we find that (1) register tokens encode useful fine-grained details that can complement the [CLS] token; (2) automatically pooled ROI tokens further improve fine-grained discrimination; and (3) multi-vector retrieval with a small set of tokens improves over a DINOv2-reg single-vector baseline while remaining tractable for large-scale search. The code is available at https://github.com/IdhcbIan/Augmenting_CLS_with_ROI_tokens.
Figures & tables
Fig. 1: Distribution of most similar tokens of the database images for any given [CLS] token from queries of images in COCO’s dog class [ 15 ] . Interestingly, a [CLS] token from a query image is often most similar to one of the four register tokens of other images (64%), not their [CLS] token (36%).
Fig. 2: For each cue token (the [CLS] and register tokens), the location of its corresponding buddy token is shown by the black max box. They focus on different parts of the semantic input: dog’s head, dog’s paw, cat’s ear, and ball.
Fig. 3: The data flow of our training setup. First, the [CLS] and register tokens are extracted with DINOv2-reg. Second, for each of these cue tokens, the buddy token is identified as the most similar image patch token. Third, the ROI token is constructed by mean pooling the buddy token and the image patch tokens in the region around it. In our experiments, we use only 10 tokens per image. Finally, the standard triplet training is performed with ColBERT-style similarity between image pairs.
Method
Dim.
Architecture
CUB-200
In-Shop
Cars-196
Stanford Online Products
1
2
4
8
1
10
20
30
1
2
4
8
1
2
4
8
NSoftmax [ 28 ]
512
ResNet50
61.3
73.9
83.5
90
86.6
97.5
98.4
98.8
84.2
90.4
94.4
96.9
78.2
90.6
96.2
-
ProxyNCA++ [ 27 ]
512
ResNet50
69.0
79.8
87.3
92.7
90.4
98.1
98.8
99.0
86.5
92.5
95.7
97.7
80.7
92.0
96.7
98.9
A-BIER [ 29 ]
512
GoogleNet
57.5
68.7
78.3
86.2
93.1
95.1
96.9
97.5
82.0
89.0
93.2
96.1
74.2
86.9
94.0
97.8
ABE [ 30 ]
512
GoogleNet
60.6
71.5
79.8
87.4
87.3
96.7
97.9
98.2
85.2
90.5
94.0
96.1
76.3
88.4
94.8
98.2
SM [ 31 ]
512
GoogleNet
56.0
68.3
78.2
86.3
90.7
97.8
98.5
98.8
83.4
89.9
93.9
96.5
75.3
87.5
93.7
97.4
TABLE I: Recall@k metrics on CUB-200, In-Shop, Cars-196, and Stanford Online Products. The table compares reported methods with different backbones, embedding dimensions, and training objectives; our DINOv2-reg ablations isolate the contribution of the proposed token selection and multi-vector scoring.
R-MAC [ 44 ]
GeM [ 45 ]
GeM+AP [ 46 ]
SSLeb [ 47 ]
43.0%
69.0%
33.5%
92.2%
SSLss [ 47 ]
SSLeb+Bd [ 47 ]
SSLss+Bd [ 47 ]
Ours
92.0%
92.5%
92.3%
93.5%
TABLE II: Comparison of retrieval performance (mAP) on the INSTRE dataset against state-of-the-art methods.
Method
CUB (R@1)
Cars-196 (R@1)
DINOv2-reg ( [CLS] only)
86.6
89.8
DINOv2-reg ( [CLS] +Registers)
86.7
90.3
DINOv2-reg ( [CLS] +Registers+ROI)
87.4
90.5
TABLE III: Effect of using register and ROI tokens (bottom 2 experiments using multi-vector training).
ROI Size
CUB (R@1)
Cars-196 (R@1)
Single Patch
87.8
89.4
3×3
88.1
90.1
5×5
88.1
90.5
7×7
88.2
90.0
9×9
87.8
90.1
TABLE IV: Ablation on the region size for ROI tokens. We report Recall@1 on CUB and Cars-196 with single patch vs. N×N mean pooling for N=3, 5, 7, 9. Our default setting is N=3.
Method
# Tokens
T. Dim
Memory
Recall@1
Single-Vector
1
384
1.5 GB
86.6
[CLS] +Registers
5
1,920
7.7 GB
86.7
Ours: [CLS] +Reg+ROI
10
3,840
15 GB
87.4
All Patch+Reg+ [CLS]
201
77,184
309 GB
87.3
TABLE V: Theoretical index size comparisons for 1M images using 384-dimensional float32 embeddings (4 bytes/dim). Recall@1 is measured on CUB-200 and does not represent million-scale retrieval performance. “T. Dim” and “Memory” denote dimensions per image and raw storage without overhead, respectively.
CUB-200
In-Shop
R@1
R@2
R@4
R@1
R@10
R@20
RRT [ 17 ]
68.7
85.0
95.6
88.3
97.9
98.6
LOCORE-tiny [ 48 ]
71.4
86.8
96.4
89.1
97.9
98.2
LOCORE-small [ 48 ]
74.6
89.1
97.3
89.4
97.7
97.7
LOCORE-base [ 48 ]
78.3
91.9
98.2
87.9
97.9
98.7
Ours
87.4
92.7
95.2
92.9
98.4
98.9
TABLE VI: Comparison of our approach with local feature-based methods on metric learning benchmarks in terms of Recall@k.
We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transformers compress each fixed-size patch into a single token, although fine-grained distinctions often depend on localized variations within only a few patches. SubViT addresses this mismatch by representing discriminative patches with multiple subtokens while retaining the original token sequence for global context, thereby allocating additional capacity where it is most needed. Since attention heads encode complementary semantics and extracting attention maps at inference requires an extra backbone forward, we adopt a two-stage training strategy. Stage 1 fine-tunes the ViT using subdivision regions sampled from random attention heads, exposing the model to diverse subdivision patterns. Stage 2 identifies informative attention maps through feature-degradation distances and distills them into a lightweight single-map router, which directly predicts deterministic token-importance scores without a separate attention forward. We evaluate SubViT on Generalized Category Discovery (GCD), a challenging task requiring both fine-grained discrimination and generalization to unlabeled novel categories. Across CUB, FGVC-Aircraft, and Stanford-Cars, SubViT improves the average novel-category accuracy of DINOv2 from 81.3% to 84.7%, with only 0.50 ms additional latency and 3.4% more FLOPs, while reducing latency by 73.8% relative to Retina Patch. Code: SubViT.
Jie Zhu, Ivy Zhang, Minchul Kim +1
Michigan State University · Cranbrook Kingswood School · University of North Carolina at Chapel Hill
Vision Transformers (ViTs) can learn strong image-level representations while their patch representations become less effective for dense prediction during prolonged training. We revisit this dense degradation phenomenon and argue that it is not fully explained by high-norm artifacts alone. Instead, we characterize \emph{semantic diffusion}: an optimization shortcut in which global semantic information spreads through patch tokens beyond what is locally justified. Our analysis shows that dense representation quality is not captured by locality alone: shallow features can remain better aligned with foreground regions yet underperform deeper features, and \texttt{[CLS]} features remain complementary for dense prediction. These observations suggest that the goal should not be to remove global context, but to make token interactions more selective. We therefore study sparse attention as a minimal intervention, replacing softmax attention with entmax-1.5 while preserving global token connectivity. On DINOv1 ViT-S/16 trained for 200 epochs on ImageNet-1K, this change preserves ImageNet linear probing accuracy and substantially improves semantic segmentation performance: VOC mIoU increases from 42.80 to 48.78, ADE20K from 19.85 to 21.97, and Cityscapes from 36.79 to 37.87. These results suggest that selective token mixing is a simple and effective bias for improving dense ViT representations.
Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.
Yaoxin Niu, Zhangquan Chen, Yang Zhang +5
Tsinghua University · Peng Cheng Laboratory · The Hong Kong University of Science and Technology +3