cs.CVSep 29, 2026

You Only Reprogram Once: Rethinking Prolonged Training for Visual Reprogramming

Authors: Zizhao Li, Mohammed Yaqoob Ansari, Xinyu Su, Jiayang Ao, Joseph West, Kourosh Khoshelham

Organizations: The University of Melbourne · Fudan University

Abstract

Visual reprogramming is a parameter-efficient method for adapting pretrained models, yet its training can remain computationally expensive: even with a frozen backbone, visual prompts are often optimized through the full model for hundreds of epochs. Before changing what the pretrained model sees, we ask whether we are fully using what it already tells us. We find that modeling the full source response can already yield strong downstream predictions without prompt optimization. Motivated by this observation, we introduce You Only Reprogram Once (YORO), which constructs a downstream predictor from the frozen response space in a single forward-only traversal. Its Bayesian Discriminant Mapping (BDM) derives a covariance-aware affine mapping from streaming class statistics, requiring no backpropagation, optimizer updates, or repeated visits to the training set. When further input adaptation helps, YORO-FP optionally refines the visual prompt for 20 epochs. BDM also extends naturally to CLIP by treating attribute-prompt similarities as source responses. Across three full-data settings, YORO improves average accuracy over the strongest prior gradient-free mapping by 18.4--24.4%. On 16-shot CLIP, it raises the four-backbone average from 71.4% to 77.2%. YORO-FP provides further gains on selected tasks, while validation often retains the one-pass predictor. These results suggest a different default for visual reprogramming: read out the frozen response first, and optimize the input only when needed.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Reprogramming Vision-Language Models via Structured Prompt Reparameterization

    Sep 29, 2026Zizhao Li, Chengyi Cai, Mohammed Yaqoob Ansari +3Contrastive Language-Image Pre-Training ModelIntra-Class

  2. ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models

    Aug 8, 2026Sajjad Ghiasvand, Yifan Yang, Mahnoosh Alizadeh +1Soft Prompt TuningFrozen Contrastive Language-Image Pre-Training Visual Encoder

  3. Low-Rank Prompt Learning for Vision-Language Models with Fixed-Token Bases

    Sep 8, 2026Tanvir Muntakim Tonoy, Sajjad Ghiasvand, Mahnoosh Alizadeh +1Multiple Prompt Learning FrameworksContrastive Language-Image Pre-Training Model