cs.CVOct 4, 2026

PCLM: Small-target localization with frozen CLIP via prototype contrast and local magnification

Authors: Zhipeng Ye, Feng Jiang, Qiufeng Wang, Hao Li

Organizations: Taizhou Institute of Science and Technology, Nanjing University of Science and Technology, Taizhou 225300, Jiangsu, China · Department of Intelligence Science, Xi’an Jiaotong-Liverpool University, Suzhou 215123, Jiangsu, China · School of Computer Science and Technology, University of Arizona, Tucson 85705, AZ, USA

Abstract

Small targets occupy few patches in a vision-language encoder, so spatial features often mix object appearance with surrounding content. We propose Prototype Contrast and Local Magnification (PCLM), a support-conditioned localization method that uses a frozen CLIP encoder. Five masked support images per class define foreground and background prototypes through equally weighted regional features. Their difference provides a shared scoring direction for query patches, explicitly comparing target evidence with the demonstrated background. Nine overlapping query windows are enlarged and encoded independently to sample small targets more densely. Reprojection and coverage averaging combine their scores into a continuous localization map. The class direction occupies 2 KiB regardless of support count and transfers unchanged across datasets with mapped categories. On 5,047 small-target queries from VOC, COCO, ADE20K and Oxford-IIIT Pets, PCLM achieves higher mean pixel AP than every evaluated text-conditioned localization baseline on each dataset under our evaluation protocol. Gains over the strongest scene-dataset baselines range from 5.63 to 13.23 percentage points. At comparable measured latency, local magnification improves scene small-target AP by 4.51 to 5.68 points over whole-canvas enlargement. Factorial experiments show that prototype contrast increases the benefit of local observation, including under matched image-coordinate filtering. Support-budget experiments show that additional examples refine category estimation without increasing representation size or query-time scoring cost.

Figures & tables

Explore similar work

CardsList
  1. Self-Improving Small Object Grounding in LVLMs

    Jun 1, 2026Tianze Yang, Yucheng Shi, Ruitong Sun +2Small ModelsObjects

  2. PosEviLoc: Position-Conditioned Spatial Evidence for Language-Based 3D Localization

    Sep 20, 2026Tianyi Shang, Yike Shi, Zhenyu LiIndoor LocalizationPoint Clouds