Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly model relationships among classes. However, fine-grained categories often exhibit highly overlapping attribute descriptions and strong inter-class correlation in the text embedding space, where discriminative cues lie in subtle low-variance components. We propose Reparameterized Inter-Class Visual Reprogramming (RVP), a structured framework that aggregates multiple text prompts within each class and applies residual correction across classes. We also show that CLIP-based visual reprogramming with input-independent linear output aggregation can be expressed as a linear mapping from frozen image embeddings to downstream logits, and use this view to design a structured reparameterization that models shared semantic components and class-specific differences. RVP uses only a single visual prompt and can be reparameterized at inference into a frozen backbone followed by a linear classifier, incurring nearly zero computational overhead. Across 11 few-shot classification benchmarks and four CLIP backbones, RVP consistently improves over prior visual reprogramming methods with comparable or better inference efficiency.
Figures & tables
Figure 1: Fine-grained classification with attribute prompts. Prompt groups from visually similar classes show high cosine similarity, indicating that attribute prompts are highly similar. The rapidly decaying singular value spectrum of the text embedding matrix further reveals strong inter-class correlation, motivating explicit modeling of class relationships.
Figure 2: Overview of Reparameterized Inter-Class Visual Reprogramming (RVP). During training, a visual prompt transforms the input image, which is then encoded by a frozen image encoder. The cosine similarities between the visual embedding and text embeddings are aggregated by a learnable intra-class matrix to produce class logits, which are further refined by a residual class-relation matrix that captures inter-class dependencies. At inference, the text embeddings and label mapping are reparameterized into a single linear classifier, enabling efficient prediction with a single forward pass.
Method
Aircraft
Caltech
Cars
DTD
ESAT
Flowers
Food
Pets
SUN
UCF
Resisc
Avg.
VP
32.1
93.5
65.5
61.4
91.2
82.5
82.3
91.0
65.8
73.8
79.1
74.4
AR
31.7
95.5
68.0
62.0
93.4
85.9
85.2
92.7
67.9
78.1
81.6
76.5
AttrVR
36.6
95.7
68.3
65.6
93.8
92.9
85.9
93.3
69.6
79.0
82.6
78.5
DVP
38.7
96.0
70.8
65.5
94.1
95.0
85.7
93.3
71.1
82.0
84.4
79.7
RVP
46.1 ± 0.2
96.5 ± 0.2
84.8 ± 0.3
68.7 ± 0.1
92.7 ± 0.2
96.7 ± 0.2
85.5 ± 0.1
94.0 ± 0.1
73.8 ± 0.1
85.1 ± 0.7
85.9 ± 0.5
82.7
Table 1: Accuracy comparison of different methods trained on 16-shot downstream classification tasks, using ViT-B/16-based CLIP as the pretrained model (Mean % ± Std %, ours are highlighted and the highest result is in bold ). Other results are taken from prior work ( Cai et al., 2025b ) .
Method
Aircraft
Caltech
Cars
DTD
ESAT
Flowers
Food
Pets
SUN
UCF
Resisc
Avg.
VP
16.2
80.1
44.0
43.4
59.7
53.6
65.3
77.2
48.8
52.0
47.7
53.5
AR
18.6
86.5
53.9
46.4
66.6
60.9
74.2
82.5
56.8
59.7
58.4
60.4
AttrVR
20.7
89.1
53.9
54.4
72.0
74.8
75.3
88.9
59.9
63.6
58.2
64.6
DVP
22.1
89.8
54.5
55.9
72.2
80.0
75.0
88.9
61.1
65.9
60.8
66.0
RVP
29.1 ± 0.3
92.0 ± 0.0
71.0 ± 0.3
62.0 ± 0.3
72.9 ± 0.9
91.6 ± 0.2
73.7 ± 0.0
89.7 ± 0.1
66.2 ± 0.1
74.3 ± 0.1
72.9 ± 0.1
72.3
Table 2: Accuracy comparison of different methods trained on 16-shot downstream classification tasks, using RN50-based CLIP as the pretrained model (Mean % ± Std %, ours are highlighted and the highest result is in bold ). Other results are taken from prior work ( Cai et al., 2025b ) .
Method
RN50
RN101
ViT-B/32
ViT-B/16
VP
53.5
57.5
68.3
74.4
AR
60.4
62.7
66.3
76.5
AttrVR
64.6
67.2
69.8
78.5
DVP
66.0
68.8
71.0
79.7
RVP
72.3
73.0
75.8
82.7
Table 3: Average accuracy of different VR methods on 11 datasets using different CLIP visual encoders (mean accuracy in %; ours are highlighted and the highest is in bold ; RN denotes ResNet).
Figure 3: Accuracy comparison across different shot settings on Aircraft using ViT-B/16 CLIP. RVP consistently outperforms prior VR methods across all shot numbers. Shaded regions indicate standard deviation.
Figure 4: Accuracy and latency comparison on FGVC Aircraft across different backbones. Our method consistently achieves the highest accuracy while maintaining low latency, demonstrating a favorable trade-off between performance and efficiency.
Method
Aircraft
Caltech
Cars
DTD
ESAT
Flowers
Food
Pets
SUN
UCF
Resisc
Avg.
RVP
46.1
96.5
84.8
68.7
92.7
96.7
85.5
94.0
73.8
85.1
85.9
82.7
w/o VR
40.1
96.2
81.7
65.7
58.5
96.8
84.7
93.9
74.2
83.6
82.8
78.0
w/o intra-class P
46.0
96.3
84.7
68.3
92.6
96.8
85.5
94.1
73.8
84.4
85.6
82.6
w/o inter-class E
35.6
96.1
68.0
63.3
93.8
91.7
85.6
93.1
67.4
78.9
83.5
77.9
Linear Probe
37.2
93.9
73.5
63.5
84.2
92.2
79.5
85.6
69.1
76.8
83.5
76.3
Attribute & LP
37.4
90.9
78.1
56.4
59.4
86.2
73.9
90.7
67.9
68.7
76.6
71.5
Table 4: Ablation studies of RVP using a ViT-B/16-based CLIP backbone. The complete method is highlighted , and the best results are shown in bold .
Figure 5: Top predicted classes and highest-matching attributes for a test image. For both RVP and DVP, the most similar attributes include prompts from other classes, reflecting strong semantic overlap in fine-grained recognition. However, RVP explicitly models inter-class relationships, allowing it to better resolve these cross-class ambiguities and produce the correct prediction.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Aircraft
Caltech
Cars
DTD
ESAT
Flowers
Food
Pets
SUN
UCF
Resisc
Task Info.
aircraft model
object
fine-grained automobile
texture
remote sensing land cover
flower
food
pet
scene
action
remote sensing scene
Class Number
100
100
196
47
10
102
101
37
397
101
45
Batch Size
64
64
64
64
64
64
64
64
64
64
64
Appendix
Table 5: Summary of the 11 downstream benchmark datasets used in our experiments, including task type, number of classes, and training batch size.
Method
Aircraft
Caltech
Cars
DTD
ESAT
Flowers
Food
Pets
SUN
UCF
Resisc
Avg.
VP
19.3
83.0
53.7
43.4
62.8
57.2
71.2
80.2
53.5
54.2
54.0
57.5
AR
19.5
89.7
62.0
46.3
70.4
60.4
78.0
84.4
58.4
60.6
60.2
62.7
AttrVR
23.3
92.0
62.2
55.6
70.3
76.2
79.5
89.3
62.1
64.5
64.5
67.2
DVP
23.8
92.7
62.5
58.0
70.7
80.6
79.1
89.5
63.7
68.1
68.4
68.8
RVP
30.4
93.9
76.0
60.1
67.6
90.5
76.8
91.1
67.5
75.8
73.4
73.0
Appendix
Table 6: Accuracy comparison of different methods trained on 16-shot downstream classification tasks, using RN101-based CLIP as the pretrained model (Mean %, ours are highlighted and the highest is in bold ).
Method
Aircraft
Caltech
Cars
DTD
ESAT
Flowers
Food
Pets
SUN
UCF
Resisc
Avg.
VP
24.3
92.3
58.6
54.9
85.9
71.2
75.0
86.8
61.0
67.3
73.9
68.3
AR
21.8
92.7
56.9
49.9
85.6
66.7
75.7
84.7
59.9
63.5
71.6
66.3
AttrVR
24.5
92.0
56.6
56.8
88.6
77.8
77.2
89.8
62.8
67.9
73.9
69.8
DVP
26.1
92.9
56.5
57.2
88.5
82.5
77.0
89.2
64.2
70.5
76.0
71.0
RVP
32.8
94.1
74.1
63.4
86.4
93.3
74.8
90.3
68.0
77.8
79.1
75.8
Appendix
Table 7: Accuracy comparison of different methods trained on 16-shot downstream classification tasks, using ViT-B/32-based CLIP as the pretrained model (Mean %, ours are highlighted and the highest is in bold ).
Method
Aircraft
Caltech
Cars
DTD
EuroSAT
Flowers
Food
Pets
SUN
UCF
RESISC
Avg.
CoOp
43.2
95.8
82.9
69.7
85.0
96.8
84.2
92.0
74.9
83.1
84.7
81.1
CoCoOp
33.3
95.1
72.3
63.7
73.6
89.1
87.4
93.4
72.6
77.2
81.6
76.3
CLIP-Adapter
34.2
94.9
74.0
59.4
71.4
92.9
87.1
92.3
74.2
80.2
85.7
76.9
Tip-Adapter-F
44.6
95.7
82.3
70.8
85.9
96.2
86.8
92.6
76.0
83.9
81.2
81.5
TaskRes
44.9
95.8
83.5
71.5
82.7
97.5
86.9
92.4
76.1
84.0
83.3
81.7
LP++
42.1
95.8
80.8
71.9
85.5
96.3
87.2
92.6
76.0
83.9
80.9
81.2
Appendix
Table 8: Accuracy comparison of different methods trained on 16-shot downstream classification tasks, using ViT-B/16-based CLIP as the pretrained model (Mean % ± Std %, ours are highlighted and the highest result is in bold ).
Method
Accuracy (%)
Trainable Params.
Inference Path
CoOp
82.9
0.008M
Fixed classifier from learned text prompts
CoCoOp
72.3
0.042M
Image-conditioned text features
CLIP-Adapter
74.0
0.131M
Nonlinear feature adapter
Tip-Adapter-F
82.3
1.606M
Cache-based adapted logits
TaskRes
83.5
0.100M
Fixed linear head
LP++
80.8
0.101M
Fixed linear head
Appendix
Table 9: Accuracy and parameter efficiency on StanfordCars under the 16-shot setting with ViT-B/16. Trainable parameters count only method-specific adaptation parameters.
Dataset
C
r90
r90/C
RVP
w/o E
ΔE
Aircraft
100
7
0.070
46.1
35.6
10.5
Caltech101
100
28
0.280
96.5
96.1
0.4
Cars
196
24
0.122
84.8
68.0
16.8
DTD
47
3
0.064
68.7
63.3
5.4
EuroSAT
10
1
0.100
92.7
93.8
-1.1
Flowers102
102
27
0.265
96.7
91.7
5.0
Appendix
Table 10: Relationship between text-space concentration and the benefit of inter-class correction. r90 denotes the minimum number of eigenvalues explaining 90% of the spectral mass of the class-prototype Gram matrix, and ΔE measures the accuracy gain from enabling the inter-class residual matrix E .
Method
Prompt Parameters
Total
Accuracy
VP
69840
69840
32.1
AR
39936
39936
31.7
AttrVR
39936
39936
36.6
DVP
119808
121808
38.7
RVP
39936
51936
46.1
Appendix
Table 11: Number of trainable parameters for different methods on the Aircraft dataset ( C=100 , M=20 ).
Aircraft
Caltech
Cars
DTD
ESAT
Flowers
Food
Pets
SUN
UCF
Resisc
Avg.
AttrVR (DesAttr)
35.9
95.6
68.2
64.4
93.8
92.4
85.7
93.0
67.7
78.6
81.8
77.9
DVP (num=1)
36.4
95.8
69.1
65.3
94.1
93.6
85.7
93.1
70.0
80.2
82.8
78.7
RVP
46.1
96.5
84.8
68.7
92.7
96.7
85.5
94.0
73.8
85.1
85.9
82.7
Appendix
Table 12: Accuracy comparison of RVP, AttrVR, and DVP under the same text prompt setting, where DVP is restricted to a single group of trainable visual prompts, using ViT-B/16 CLIP as the pretrained model (mean %; ours are highlighted and the best results are shown in bold ).
Aircraft
Caltech
Cars
DTD
ESAT
Flowers
Food
Pets
SUN
UCF
Resisc
Avg.
AttrVR
36.6
95.7
68.3
65.6
93.8
92.9
85.9
93.3
69.6
79.0
82.6
78.5
DVPlite
39.3
95.9
71.4
66.5
93.8
95.2
85.8
93.4
71.6
81.0
83.6
79.8
RVP
46.1
96.5
84.8
68.7
92.7
96.7
85.5
94.0
73.8
85.1
85.9
82.7
Appendix
Table 13: Accuracy comparison of our RVP and DVPlite trained on 16-shot downstream classification task, using ViT-B/16-based CLIP as the pretrained model (Mean %, ours is highlighted and the highest is in bold ).
Figure 6: Additional visualization of the top predicted classes and highest-matching attributes for a test image. For both RVP and DVP, the most similar attributes include prompts from other classes, reflecting strong semantic overlap in fine-grained recognition. However, RVP explicitly models inter-class relationships, allowing it to better resolve these cross-class ambiguities and produce the correct prediction.
Figure 7: Qualitative examples on Food101. Food categories often exhibit large intra-class variation and strong cross-class visual overlap due to differences in plating, viewpoint, garnish, and accompanying side dishes. As shown here, classes such as Apple Pie and Waffles can be confused when the main dish is partially visible or co-occurs with similar desserts, while Beet Salad and Tuna Tartare may share similar fine-grained presentation and ingredients. Correct predictions are shown in green and incorrect predictions in red.
Figure 8: Qualitative examples on EuroSAT. Several classes, especially Sea_or_lake, River, and Highway/Road, exhibit strong visual similarity in satellite crops due to elongated shapes, curved boundaries, and limited scene context. As a result, some samples remain ambiguous even under the proposed method, suggesting that EuroSAT is less dominated by inter-class semantic ambiguity than fine-grained recognition benchmarks. Correct predictions are shown in green and incorrect predictions in red.
Figure 9: Memory consumption when training RVP with a ViT-B/16-based CLIP backbone on the FGVC Aircraft dataset. Left: host memory. Right: GPU memory.
Abbreviation
Description
RVP
Reparameterized Inter-Class Visual Reprogramming.
VR
Visual Reprogramming.
VLM
Vision-Language Model.
CLIP
Contrastive Language-Image Pre-training.
VP
Visual Prompting / standard visual reprogramming baseline.
AR
Adversarial Reprogramming baseline.
Appendix
Table 14: Abbreviations used in the paper
Symbol
Description
fimg
CLIP image encoder.
ftxt
CLIP text encoder.
XS
Source image space of the pretrained CLIP model, with XS⊆RdS .
XT
Target image space for the downstream task, with XT⊆RdT .
YT
Label space of the downstream task, with YT={1,…,C} .
V
Text space containing textual descriptions.
Appendix
Table 15: Generic notation in visual reprogramming
Symbol
Description
T
Stacked matrix of normalized text embeddings. In RVP with C classes and M descriptions per class, T∈RCM×D .
Ma
Vector of similarity scores over all textual descriptions.
My
Vector of downstream class logits.
P∈RC×M
Learnable intra-class weighting matrix for aggregating attribute descriptions within each class.
P~c
Softmax-normalized weight vector for the c -th class, obtained from the c -th row of P .
P~c,m
Normalized weight assigned to the m -th textual description of class c .
Appendix
Table 16: Notation for visual reprogramming and the linear mapping view
KT Corporation, Republic of Korea · Pohang University of Science and Technology (POSTECH), Republic of Korea · National AI Research Lab, Republic of Korea