PCLM: Small-target localization with frozen CLIP via prototype contrast and local magnification
Authors: Zhipeng Ye, Feng Jiang, Qiufeng Wang, Hao Li
Organizations: Taizhou Institute of Science and Technology, Nanjing University of Science and Technology, Taizhou 225300, Jiangsu, China · Department of Intelligence Science, Xi’an Jiaotong-Liverpool University, Suzhou 215123, Jiangsu, China · School of Computer Science and Technology, University of Arizona, Tucson 85705, AZ, USA
Small targets occupy few patches in a vision-language encoder, so spatial features often mix object appearance with surrounding content. We propose Prototype Contrast and Local Magnification (PCLM), a support-conditioned localization method that uses a frozen CLIP encoder. Five masked support images per class define foreground and background prototypes through equally weighted regional features. Their difference provides a shared scoring direction for query patches, explicitly comparing target evidence with the demonstrated background. Nine overlapping query windows are enlarged and encoded independently to sample small targets more densely. Reprojection and coverage averaging combine their scores into a continuous localization map. The class direction occupies 2 KiB regardless of support count and transfers unchanged across datasets with mapped categories. On 5,047 small-target queries from VOC, COCO, ADE20K and Oxford-IIIT Pets, PCLM achieves higher mean pixel AP than every evaluated text-conditioned localization baseline on each dataset under our evaluation protocol. Gains over the strongest scene-dataset baselines range from 5.63 to 13.23 percentage points. At comparable measured latency, local magnification improves scene small-target AP by 4.51 to 5.68 points over whole-canvas enlargement. Factorial experiments show that prototype contrast increases the benefit of local observation, including under matched image-coordinate filtering. Support-budget experiments show that additional examples refine category estimation without increasing representation size or query-time scoring cost.
Figures & tables
Figure 1: A small-boat example from VOC. CCI [ 2 ] uses class text and one global view (17.0% pixel AP); PCLM uses five masked supports from bank 11 and nine local views (96.5% AP). The target occupies 2.51% of the canvas. AP uses original scores over known foreground and background pixels; each heatmap is scaled separately for display.
Method (venue/year)
Target
Spatial computation
Heatmap score
CCI (CVPR’26) [ 2 ]
Class text
Final-block cluster masking
Normalized cosine-similarity drop
Grad-ECLIP (TPAMI’26) [ 8 ]
Class text
Channel and spatial weighting
Matching relevance
CLIP Surgery (PR’25) [ 11 ]
Class text
Attention and feature surgery
Dense image–text similarity
Grad-CAM (IJCV’20) [ 6 ]
Class text
Activation–gradient aggregation
Positive class activation
gScoreCAM (ACCV’22) [ 10 ]
Class text
Gradient-selected mask scoring
Score-weighted activation
LeGrad (ICCV’25) [ 9 ]
Class text
Positive attention gradients
Feature-formation sensitivity
Table 1: Target information and spatial readouts in the localization comparison. Parentheses give publication venues and years. Scores describe the evaluated outputs.
Figure 2: PCLM. (a) Five support images and their masks define foreground and background prototypes, with equal weight for each image. Their difference gives the reusable class direction dc . (b) Nine overlapping query windows are enlarged and independently encoded. The shared direction produces signed patch scores, followed by mean filtering, nearest-neighbor reprojection and coverage averaging. Local magnification halves the nominal patch footprint on the query canvas and allocates approximately four times as many patch areas to a visible target. The boat example illustrates the observation and fusion stages.
Dataset
Images
Query pairs
Categories
Small pairs
VOC validation
1,449
2,147
20
655
COCO validation
4,031
6,900
20
2,978
ADE20K validation
1,218
2,015
12
1,338
Pets test
3,662
3,662
2
76
Table 2: Evaluation queries after fixed label conversion and eligibility checks. Small denotes category-union foreground below 5% of the full 224×224 canvas. All supports are from VOC training.
Approach and observation
VOC
COCO
ADE20K
Pets
CCI, global
11.57
5.61
3.71
47.67
Grad-CAM, global
23.84
18.90
19.13
13.50
Grad-ECLIP, global
35.88
22.46
20.88
73.05
Grad-ECLIP + nine
45.84
29.12
26.96
86.26
CLIP Surgery, global
37.06
26.01
22.41
63.20
CLIP Surgery + nine
50.76
36.08
30.63
81.72
Table 3: Small-target AP (%) for text-conditioned localization adapters and PCLM. The text block receives class names; PCLM receives five masked supports. “+ nine” denotes the nine-window extension.
Figure 3: Selected successful small-target localizations across four datasets. PCLM uses five supports from bank 11 and nine views; baseline view counts are shown above each column. CCI uses final-block masking and signed cluster-score normalization. Green denotes foreground, light gray background and hatched gray ignored pixels. Per-map 1st–99th percentile scaling affects display only; AP uses scores before display scaling.
Figure 4: Selected successful ADE20K and Pets examples. PCLM uses five masked supports from bank 11; the references use class text. All local maps use nine windows, retaining each method’s per-view scores before projection and coverage averaging. Green contours mark targets and gray shading marks ignored pixels. Display uses per-map 1st–99th percentile scaling; AP uses unscaled scores over known pixels.
Query set
G
4
9
Gain
95% CI
VOC, all
73.14
78.76
81.35
8.21
[7.61,8.82]
VOC, small
44.25
60.07
63.99
19.73
[18.38,21.15]
COCO, all
54.51
60.15
62.60
8.09
[7.77,8.40]
COCO, small
27.70
39.07
41.71
14.00
[13.45,14.56]
ADE20K, all
38.78
46.84
48.71
9.93
[9.30,10.57]
ADE20K, small
24.34
34.93
36.86
12.52
[11.71,13.34]
Table 4: PCLM localization with fixed five-shot supports. AP is percent. G/4/9 denote global 224, four corners and nine windows. Gain and paired 95% interval are nine minus global in percentage points. Means average five banks.
Figure 5: Who benefits from local magnification? Left: empirical cumulative distributions of nine-minus-global AP changes after averaging five support banks. Right: scene-dataset mean gains by target-area bin with paired parent-image 95% intervals. Area bins are <1% , [1%,2%) and [2%,5%) of the canvas.
Observation and filter
VOC
COCO
ADE20K
Pets
Global 224, patch 3
44.25
27.70
24.34
88.58
Four windows, patch 3
60.07
39.07
34.93
95.14
Nine windows, patch 3
63.99
41.71
36.86
97.01
Global 224, canvas 25
59.30
38.92
34.12
94.88
Nine windows, canvas 25
66.39
43.52
38.34
97.73
Table 5: View layouts and common-coordinate filtering with fixed PCLM prototypes: small-target AP (%). Patch 3 applies a 3×3 filter on each patch map; Canvas 25 applies a 25×25 filter in output coordinates.
(a) Resolution, image source and filtering
Input
Source
Filter
VOC
COCO
ADE20K
Pets
448
Canvas
None
56.47 / 62.26
36.44 / 38.19
32.71 / 31.91
88.41 / 91.05
448
Canvas
Patch 3
59.89 / 71.42
37.57 / 45.94
33.81 / 40.01
96.37 / 99.74
448
Canvas
Canvas 25
61.51 / 71.48
38.66 / 46.82
34.79 / 40.22
97.21 / 100.00
448
Original
None
57.26 / 64.03
37.91 / 40.20
33.69 / 33.24
88.36 / 93.42
448
Original
Patch 3
61.09 / 72.82
39.34 / 47.90
35.08 / 41.49
96.87 / 100.00
Table 6: Global resolution, input source and spatial filtering. (a) Small-target AP / first-maximum Pointing (%). (b) Nine-view minus global paired AP gain and 95% interval under patch-3 filtering, in percentage points. All five banks and all 5,047 queries are used. Canvas inputs enlarge the shared 224-pixel image; Original RGB inputs resize the source image directly. Global resolution is an input side length; 9 views use nine independently encoded 224×224 forwards.
Scoring / fusion
VOC
COCO
ADE20K
Pets
Foreground similarity only
36.20
19.49
16.13
84.90
Foreground–background raw margin
63.99
41.71
36.86
97.01
Per-view min–max margin
46.16
27.10
24.99
91.18
Per-view standardized margin
56.73
36.72
32.76
92.27
sigmoid(20s)−0.5
64.03
41.87
37.05
97.11
Table 7: Main-method controls: nine-view small-target AP (%). Foreground-only uses the exact PCLM foreground center. All score transforms are applied before patch-3 smoothing.
Dataset
No smoothing
Patch 3×3
Canvas 25×25
VOC
17.58 [16.54,18.65]
15.02 [13.64,16.45]
3.45 [1.99,4.89]
COCO
12.36 [11.94,12.79]
11.91 [11.37,12.45]
2.96 [2.44,3.49]
ADE20K
10.85 [10.18,11.51]
11.52 [10.73,12.31]
3.35 [2.58,4.11]
Table 8: Interaction between foreground–background contrast and local observation on small scene targets (AP percentage points). The interaction is (P9−PG)−(F9−FG) , where P is PCLM and F uses its foreground center alone. Each cell gives the interaction and paired parent-image 95% interval. Canvas 25 applies the same image-coordinate filter to global and local maps.
Figure 6: Nine-view localization over nested support budgets. Left: all-query AP; right: small-target AP. Shading shows one sample standard deviation over five support banks. The stored class representation has the same size at every support budget.
Layout
Views
ms
VOC
COCO
ADE20K
Pets
Global
1
7.27
73.14 / 44.25
54.51 / 27.70
38.78 / 24.34
97.91 / 88.58
Four corners
4
14.78
78.76 / 60.07
60.15 / 39.07
46.84 / 34.93
98.86 / 95.14
Corners + center
5
16.74
80.30 / 62.58
61.50 / 40.59
47.78 / 35.90
99.08 / 96.59
Corners + top/bottom
6
19.75
79.96 / 62.11
61.24 / 40.22
47.51 / 35.57
99.12 / 96.36
Corners + left/right
6
19.94
80.29 / 62.28
61.54 / 40.65
48.00 / 36.15
99.10 / 96.25
Nine windows
9
27.36
81.35 / 63.99
62.60 / 41.71
48.71 / 36.86
99.27 / 97.01
Table 9: Fixed PCLM view layouts, complete-query AP and measured inference time. Each dataset cell is overall / small-target AP (%). All crop layouts retain the four corners. Time covers in-memory RGB to CPU heatmap on RTX 4090, excluding support preparation and disk decoding; every row belongs to the same timing series.
Dataset
Views
AP
Small AP
Gain
95% CI
VOC
Global
74.42
49.76
—
—
VOC
Four corners
78.83
64.91
—
—
VOC
Nine windows
81.46
68.82
7.04
[6.45,7.64]
ADE20K
Global
41.22
27.92
—
—
ADE20K
Four corners
48.07
37.94
—
—
ADE20K
Nine windows
49.97
39.90
8.75
[8.12,9.38]
Table 10: PCLM on ViT-L/14 block 18 with unchanged VOC support identities. AP is percent; gain and interval are nine-window minus global overall AP in percentage points. Every valid VOC and ADE20K query is evaluated.
Annotation
VOC PCLM
COCO PCLM
Clean
80.32
66.58
Erode 1
80.29
66.55
Erode 3
80.12
66.26
Dilate 1
80.32
66.58
Dilate 3
80.24
66.51
Table 11: Support-boundary diagnostic on 400 class-balanced queries per dataset (AP, %), averaged over five support banks. Radii are canvas pixels; erosion retains a foreground anchor. All paths use the fixed nine windows.
Figure 7: Selected failures of local magnification. Global and nine-window PCLM use the same five supports from bank 11. Green contours mark the target and gray shading marks ignored pixels. Heatmaps use per-map 1st–99th percentile display scaling; AP uses original scores over known pixels. Cases were selected from bank-11 small-target queries with global AP of at least 40%, nine-window AP of at most 35% and an AP decrease of at least 20 percentage points, followed by visual inspection for one identifiable target per scene dataset.