cs.CVOct 2, 2026

LAS-CLIP: A Lightweight Adapter Steering Approach for CLIP's Visual Encoder

Authors: Anh-Khoa Dinh-Duc, Duc-Tai Dinh, Tam V. Nguyen, Minh-Triet Tran

Organizations: University of Science, Ho Chi Minh City, Vietnam · Viet Nam National University, Ho Chi Minh City, Vietnam · University of Dayton, USA

Abstract

CLIP's visual encoder produces only global image representations, limiting its use in region-level tasks. Existing adaptations rely on visual prompting, input masking, or encoder fine-tuning, each compromising pre-trained representations. We propose LAS-CLIP, a Lightweight Adapter Steering approach that keeps every CLIP parameter frozen. A compact MaskAdapter generates per-head, per-layer attention biases from an input mask and injects them into the frozen self-attention layers, steering attention toward the target region. Crucially, because the backbone remains strictly untouched, LAS-CLIP seamlessly reverts to vanilla CLIP when no mask is provided, preserving its foundational zero-shot capabilities. With approximately 116K to 145K trainable parameters and 100K training samples on two T4 GPUs, LAS-CLIP achieves competitive or superior results compared to Alpha-CLIP on ImageNet-S zero-shot classification and RefCOCO referring expression comprehension, despite the latter fine-tuning its entire encoder on millions of samples. Qualitative analysis further confirms stronger representational fidelity under incorrect masks and in downstream generation.

Figures & tables

Explore similar work

CardsList
  1. Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models

    May 15, 2026Fabian Morelli, Arnas Uselis, Ankit Sonthalia +1Contrastive Language-Image Pre-Training ModelModel Fine-Tuning

  2. PERL: Parameter Efficient Reasoning in CLIP Latent Space

    May 18, 2026Simone Carnemolla, Salvatore Calcagno, Daniela Giordano +2Vision-Language Model AdaptationContrastive Language-Image Pre-Training Model

  3. Improving CLIP Adaptation by Breaking Tail Alignment for Source-Free Cross-Domain Few-Shot Learning

    May 28, 2026Shuai Yi, Yixiong Zou, Yuhua Li +1Vision-Language Model AdaptationContrastive Language-Image Pre-Training Model