cs.CVJan 22, 2025

TeD-Loc: Text Distillation for Weakly Supervised Object Localization

Authors: Shakeeb Murtaza, Soufiane Belharbi, Alexis Guichemerre, Marco Pedersoli, Eric Granger

Organizations: LIVIA, ILLS, Dept. of Systems Engineering, ETS Montreal, Canada

Abstract

Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, without requiring explicit bounding-box annotations. Given their reliance on classification objectives, traditional WSOL methods, like class activation mapping, tend to focus on the most discriminative object regions, often missing the full spatial extent. Although vision-language models like CLIP encode rich semantic priors, their global text and class-token embeddings are not explicitly aligned with local patch embeddings, limiting patch-level localization. Recent methods such as GenPrompt address this limitation, but at the cost of increased complexity, as they rely on conditional denoising and elaborate prompt-learning strategies. In this paper, we propose Text Distillation for Localization (TeD-Loc), which distills knowledge from CLIP text embeddings to patch embeddings through contrastive alignment, thereby enabling patch-level foreground/background localization. A localization-guided classification module is also introduced, which uses localization scores to aggregate foreground patch embeddings for joint classification and localization within a single model. In addition, a QR-based orthogonalization of class text embeddings is applied before distillation to improve discrimination for semantically similar classes. Extensive experiments show that TeD-Loc improves Top-1 Loc by ~5% on CUB and ILSVRC, and PxAP by ~31% on histopathology benchmarks, while achieving more efficient inference than GenPrompt.

Figures & tables

Explore similar work

CardsList
  1. A Realistic Protocol for Evaluation of Weakly Supervised Object Localization

    Apr 15, 2024Shakeeb Murtaza, Soufiane Belharbi, Marco Pedersoli +1Bounding BoxLocalization

  2. What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization

    May 12, 2026Ryota Yoshihashi, Masahiro Kada, Satoshi Ikehata +2Vision TransformerObject Localization

  3. Mechanisms of Object Localization in Vision-Language Models

    May 19, 2026Timothy Schaumlöffel, Martina G. Vilas, Gemma RoigObject LocalizationLocalization