Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, without requiring explicit bounding-box annotations. Given their reliance on classification objectives, traditional WSOL methods, like class activation mapping, tend to focus on the most discriminative object regions, often missing the full spatial extent. Although vision-language models like CLIP encode rich semantic priors, their global text and class-token embeddings are not explicitly aligned with local patch embeddings, limiting patch-level localization. Recent methods such as GenPrompt address this limitation, but at the cost of increased complexity, as they rely on conditional denoising and elaborate prompt-learning strategies. In this paper, we propose Text Distillation for Localization (TeD-Loc), which distills knowledge from CLIP text embeddings to patch embeddings through contrastive alignment, thereby enabling patch-level foreground/background localization. A localization-guided classification module is also introduced, which uses localization scores to aggregate foreground patch embeddings for joint classification and localization within a single model. In addition, a QR-based orthogonalization of class text embeddings is applied before distillation to improve discrimination for semantically similar classes. Extensive experiments show that TeD-Loc improves Top-1 Loc by ~5% on CUB and ILSVRC, and PxAP by ~31% on histopathology benchmarks, while achieving more efficient inference than GenPrompt.
Figures & tables
Fig. 1: A comparison of our TeD-Loc versus CLIP-ES [ 25 ] methods for extracting localization maps from CLIP at test time. (A) CLIP-ES utilizes Grad-CAM to extract localization maps from CLIP, requiring GT class labels during inference. (B) In contrast, our TeD-Loc model distills knowledge from CLIP text embeddings into the visual encoder during training, allowing it to produce both classification scores and localization maps without requiring class labels during inference.
Fig. 2: An overview of the TeD-Loc method for distilling FG text embeddings into the patch embedding backbone. First, pseudo-labels are extracted to guide the identification of FG and BG patches. By leveraging these FG/BG regions, the model minimizes the similarity of EV with the relevant text embedding for FG classes while maximizing dissimilarity with embeddings of other classes. Using a binary FG/BG classifier, TeD-Loc generates localization maps by classifying patches as FG or BG and producing class probabilities for image classification. This joint task enables the model to produce both accurate localization and classification outputs without explicit bounding box supervision.
Fig. 3: t-SNE visualizations of CLIP text embeddings for ILSVRC [ 13 ] classes before and after orthogonalization. (Left) Prior to orthogonalization, embeddings of semantically similar classes (e.g., “airplane” and “aircraft”) cluster closely together, leading to potential confusion. (Right) After orthogonalization (QR decomposition), the embeddings are more uniformly distributed and orthogonal, reducing overlap.
GlaS
C17-0
C17-1
C17-2
C17-3
C17-4
Train
67
634
1066
498
816
940
Val (CL)
18
144
38
110
102
146
Val (PxAP)
6
10
10
10
10
10
Test
80
262
172
448
376
298
Table 1: Dataset splits used in [ 19 ] for CAMELYON17 .
Method
ILSVRC
CUB
MaxBoxAcc
Top-1 Loc
Top-5 Loc
MaxBoxAcc
Top-1 Loc
Top-5 Loc
CLIP-ES (Pred) [ 25 ] (cvpr,2023)
71.2
–
–
91.6
–
–
TS-CAM [ 16 ] (iccv,2021)
67.7
53.4
64.3
71.3
83.8
87.7
SCM [ 1 ] (eccv,2022)
68.8
56.1
66.4
76.4
91.6
96.6
LCTR [ 8 ] (aaai,2022)
68.7
56.1
65.8
92.4
79.2
89.9
PSOL [ 48 ] (cvpr,2020)
66.3
58.0
65.0
91.8
80.9
90.0
Table 2: MaxBoxAcc , Top-1 Loc , and Top-5 Loc performance of TeD-Loc against state-of-the-art methods on the ILSVRC and CUB datasets. The first row corresponds to Grad-CAM for CLIP [ 25 ] , where class labels are required to produce text embeddings for extracting localization maps. Noteably, TeD-Loc outperforms existing methods without relying on text-encoder. The row TeD-Loc * ( ⟨ patch, anchor ⟩ ) reports a GT-known analysis variant, where localization maps are computed from patch–class similarity scores via the dot product ⟨zp,ty⟩ . Here, zp denotes the patch embedding at location p∈Ω , and ty is the text anchor associated with the ground-truth class y . This variant is used to assess the quality of the text-to-patch alignment.
Methods
Complexity Analysis
# Parameters
Inference Time
GenPrompt [ 51 ]
1133.35M
272ms
TeD-Loc (ours)
569.67M
121ms
Table 3: Computational complexity of our proposed TeD-Loc against GenPrompt.
Fig. 4: Qualitative comparison on ILSVRC data. Localization maps are obtained via (patch, class) embeddings dot product: ⟨zp,ty⟩ where zp,∀p∈Ω are the patch embeddings. GenPrompt fails to localize objects in complex scenes, often because it relies on external classifiers to compute text embeddings. This can fail if the classifier makes mistakes. This dependence on class labels during inference highlights GenPrompt’s vulnerability to localization errors. In contrast, TeD-Loc can localize objects in complex scenes. Here, green bboxes denote GT localization, while red bboxes represent predicted localizations.
Methods
GlaS
C17-0
C17-1
C17-2
C17-3
C17-4
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
DeepMIL [ 22 ] (icml,2018)
79.9
100.0
34.8
82.8
30.1
80.8
31.3
63.8
27.6
88.6
18.0
59.4
GradCAM ++ [ 6 ] (wacv,2018)
76.8
100.0
21.9
72.1
22.2
66.3
29.8
80.8
32.4
77.6
21.4
59.7
LayerCAM [ 23 ] (tip,2021)
75.1
100.0
22.8
72.1
22.6
66.3
30.1
80.8
33.1
77.6
21.8
66.1
SAT [ 40 ] (iccv,2023)
65.9
100.0
20.6
64.5
17.5
72.7
27.7
65.6
20.4
52.6
18.9
49.6
PixelCAM [ 18 ] (midl,2025)
86.6
100.0
49.8
80.9
52.2
73.8
54.9
73.4
71.9
65.1
50.6
61.1
Table 4: Localization ( PxAP ) and classification ( CL ) accuracy on the GlaS and CAMELYON17 center-wise test datasets.
Methods
GlaS
C17-0
C17-1
C17-2
C17-3
C17-4
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
PxAP ( ↑ )
CL ( ↑ )
TeD-Loc wo/ CONCH [ 26 ]
49.3
53.7
16.1
50.0
14.7
50.0
27.2
50.0
20.9
50.0
19.2
50.0
TeD-Loc w/ CONCH [ 26 ]
53.5
53.7
16.7
50.0
14.8
50.0
28.0
50.0
19.8
50.0
19.4
50.0
Table 5: Zero-shot localization ( PxAP ) and classification ( CL ) accuracies of TeD-Loc with and without CONCH initialization on GlaS and CAMELYON17 center-wise test sets.
Fig. 5: Visualization of localization map obtained via TeD-Loc as compared to standard WSOL models on the C17-4 dataset for the class cancer.
Fig. 6: Visualization of localization map obtained via TeD-Loc on the C17-4 dataset for the class cancer with different CAM modules (DeepMil [ 22 ] , GradCAM++ [ 6 ] , LayerCAM [ 23 ] , PixelCAM [ 18 ] ).
Losses
CUB ( MaxBoxAcc )
LKD
62.3
LKD + PCL
95.4
LKD + PCL + ICL
98.7
Table 6: Ablation study on the CUB dataset showing the impact of different loss combinations on MaxBoxAcc performance.
Dataset
λ1
λ2
λ3
0.5
1.0
1.5
2.0
0.5
1.0
1.5
2.0
0.5
1.0
1.5
2.0
GlaS
89.1
88.8
78.7
78.2
78.1
88.8
78.6
78.6
78.3
88.8
86.4
78.8
Table 7: Ablation study of the loss weights on GlaS . Columns report the values tested for λ1 , λ2 , and λ3 .
Methods
C17-4
PxAP ( ↑ )
CL ( ↑ )
DeepMIL [ 22 ]
18.0
59.4
TeD-Loc w/ DeepMIL
78.3
89.3
GradCAM ++ [ 6 ]
21.4
59.7
TeD-Loc w/ GradCAM ++
81.1
94.3
LayerCAM [ 23 ]
21.8
66.1
Table 8: Impact of different CAM-based pseudo-labels on TeD-Loc performance. Localization ( PxAP ) and classification ( CL ) accuracies on test set of C17-4 dataset.
Fig. 7: Visualization of localization map obtained via g(zp) as compared to (patch, class) embeddings dot product: ⟨zp,ty⟩ where zp,∀p∈Ω is the patch embeddings, and y is the true image class over different variants of CLIP , and our method. TE@PE is the vanilla CLIP [ 32 ] where TE is the text embedding, and PE is the patch embedding.
Fig. 8: Failure cases of TeD-Loc in localizing ROIs. Row (a) shows 3 examples from natural images ( ILSVRC ) where GT annotations are bboxes and localization is evaluated via MaxBoxAcc . Row (b) shows 3 examples from histology images ( CAMELYON17 ), where GT annotations are pixel-level segmentation masks, and localization is evaluated via PxAP
CUB
ILSVRC
MaxBoxAcc
Top-1 Loc
Top-5 Loc
MaxBoxAcc
Top-1 Loc
Top-5 Loc
Vanilla CLIP [ 32 ]
18.8
8.9
15.2
41.1
26.6
37.0
SCLIP (eccv’24) [ 38 ]
85.8
14.4
37.1
70.4
33.9
55.1
NACLIP (wacv’25) [ 20 ]
80.8
12.0
32.3
71.7
28.6
49.4
TeD-Loc w/ g(zp)
75.6
70.0
75.1
98.7
91.7
97.6
TeD-Loc w/ ⟨zp,ty⟩
77.2
71.8
76.7
98.7
92.0
97.5
Table 9: Localization performance of maps obtained via g(zp) and (patch, class) embeddings dot product: ⟨zp,ty⟩ where zp,∀p∈Ω is the patch embeddings, and y is the true image class. We report localization performance ( MaxBoxAcc , Top-1 Loc , Top-5 Loc ) using different variants of CLIP and our method.
TeD-Loc
CUB
w/o orthogonalization, with default anchors and g(zp)
97.7
56.0
54.8
85.8
w/ orthogonalization, with g(zp)
98.7
93.0
91.7
97.6
Table 10: Impact of class text embeddings (text anchors) orthogonalization over localization and classification performance in our method over CUB dataset.
Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-level labels. The absence of bounding box (bbox) supervision during training raises challenges in the literature for hyper-parameter tuning, model selection, and evaluation. WSOL methods rely on a validation set with bbox annotations for model selection, and a test set with bbox annotations for threshold estimation for producing bboxes from localization maps. This approach, however, is not aligned with the WSOL setting as these annotations are typically unavailable in real-world scenarios. Our initial empirical analysis shows a significant decline in LOC performance when model selection and threshold estimation rely solely on class labels and the image itself, respectively, compared to using manual bbox annotations. This highlights the importance of incorporating bbox labels for optimal model performance. In this paper,a new WSOL evaluation protocol is proposed that provides LOC information without the need for manual bbox annotations. In particular, we generated noisy pseudo-boxes from a pretrained off-the-shelf region proposal method such as Selective Search, CLIP, and RPN for model selection. These bboxes are also employed to estimate the threshold from LOC maps, circumventing the need for test-set bbox annotations. Our experiments with several WSOL methods on challenging natural and medical image datasets show that using the proposed pseudo-bboxes for validation facilitates the model selection and threshold estimation, with LOC performance comparable to models selected using GT bboxes on the validation set and threshold estimation on the test set. It also outperforms models selected using class-level labels, and then dynamically thresholded with only LOC maps.
Shakeeb Murtaza, Soufiane Belharbi, Marco Pedersoli +1
LIVIA, ILLS, Dept. of Systems Engineering, ETS Montreal, Canada
Many image understanding tasks involve identifying what is present and where it appears. However, tasks that address where, such as object discovery, detection, and segmentation, are often considerably more complex than image classification, which primarily focuses on what. One possible reason is that classification-oriented backbones tend to emphasize semantic information about what, while implicitly entangling or suppressing information about where. In this work, we focus on an inductive bias termed what-where separation, which encourages models to represent object appearance and spatial location in a decomposed manner. To incorporate this bias throughout an attentive backbone in the style of Vision Transformer (ViT), we propose the What-Where Transformer (WWT). Our method introduces two key novel designs: (1) it treats tokens as representations of what and attention maps as representations of where, and processes them in concurrent feed-forward modules via a multi-stream, slot-based architecture; (2) it reuses both the final-layer tokens and attention maps for downstream tasks, and directly exposes them to gradients derived from task losses, thereby facilitating more effective and explicit learning of localization. We demonstrate that even under standard single-label classification-based supervision on ImageNet, WWT exhibits emergent multiple object discovery directly from raw attention maps, rather than via additional postprocessing such as token clustering. Furthermore, WWT achieves superior performance compared to ViT-based methods on zero-shot object discovery and weakly supervised semantic segmentation, and it is transferable to various localization setups with minimal modifications. Code will be published after acceptance.
Visually-grounded language models (VLMs) are highly effective in linking visual and textual information, yet they often struggle with basic classification and localization tasks. While classification mechanisms have been studied more extensively, the processes that support object localization remain poorly understood. In this work, we investigate two representative families, LLaVA-1.5 and InternVL-3.5, using a suite of mechanistic interpretability tools, including token ablations, attention knockout, and causal mediation analysis. We find that localization is driven by a containerization mechanism in which object-aligned tokens define the spatial extent of the object, while the semantic arrangement of tokens within those boundaries is largely irrelevant to the predicted box. Only a very small set of attention heads mediates the causal effect for both classification and localization, concentrating in early-mid layers for LLaVA and mid-late layers for InternVL. The two tasks share some early processing but ultimately depend on largely distinct specialized heads. Overall, we provide the first layer- and head-level account of localization in VLMs, revealing narrow computational pathways that can guide future model design and grounding objectives.
Timothy Schaumlöffel, Martina G. Vilas, Gemma Roig
Goethe University Frankfurt, Germany · The Hessian Center for AI, Germany