cs.CVOct 8, 2026

DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning

Authors: Wenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han, Yilong Yin, Liqiang Nie

Organizations: School of Software, Shandong University, Jinan 250101, China · Shenzhen Loop Area Institute, Shenzhen 518038, China · School of Computer and Artificial Intelligence, Shandong University of Finance and Economics, Jinan 250014, China · School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), 518055, China

Abstract

Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making support prototypes susceptible to contextual contamination. To address this problem, we propose DVLA-RL++, which extends DVLA-RL with complementary semantic purification (CSP) and counterfactual reinforcement-learning gating (CRG). Specifically, CSP generates intrinsic and nuisance descriptions from labeled supports and compares their agreement with each support token. An ambiguity-dependent rejection margin guides sparse evidence allocation, while an intrinsic semantic anchor fills the unassigned mass to provide a fallback when visual evidence is unreliable. CRG learns layer-wise semantic fusion strengths using a reward that balances recognition performance and nuisance exposure. An independently executed reference trajectory on the same episode provides a paired learning signal. Theoretical analysis relates retained evidence and anchor quality to prototype stability and establishes conditions for unbiased on-policy gradient estimation. Experiments on standard, fine-grained, and cross-domain benchmarks show state-of-the-art accuracy, with an average gain of 1.4% over DVLA-RL. The project page is available at https://peacelwh.github.io/TPAMI27-DVLA-RLpp/.

Figures & tables

Explore similar work

CardsList
  1. Project and Mix: Task-Semantic Prototypes for Few-Shot Image Classification

    Mar 25, 2026Dipam Goswami, Simone Magistri, Gido M. van de Ven +4Prototype-Based LearningVLM Adaptation

  2. PRiSM: Prototype Regularization for Few-Shot VLMs

    Jul 20, 2026Ghassen Baklouti, Omprakash Chakraborty, Jose Dolz +1VLM AdaptationPrototype Learning

  3. Rethinking Prototype-based Similarity Learning for Few-Shot Object Detection

    Jun 22, 2026KunHo Heo, Seungjae Kim, Wongyu Lee +2Prototype-Based ClassificationFew-Shot Object Detection