Organizations: School of Software, Shandong University, Jinan 250101, China · Shenzhen Loop Area Institute, Shenzhen 518038, China · School of Computer and Artificial Intelligence, Shandong University of Finance and Economics, Jinan 250014, China · School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), 518055, China
Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making support prototypes susceptible to contextual contamination. To address this problem, we propose DVLA-RL++, which extends DVLA-RL with complementary semantic purification (CSP) and counterfactual reinforcement-learning gating (CRG). Specifically, CSP generates intrinsic and nuisance descriptions from labeled supports and compares their agreement with each support token. An ambiguity-dependent rejection margin guides sparse evidence allocation, while an intrinsic semantic anchor fills the unassigned mass to provide a fallback when visual evidence is unreliable. CRG learns layer-wise semantic fusion strengths using a reward that balances recognition performance and nuisance exposure. An independently executed reference trajectory on the same episode provides a paired learning signal. Theoretical analysis relates retained evidence and anchor quality to prototype stability and establishes conditions for unbiased on-policy gradient estimation. Experiments on standard, fine-grained, and cross-domain benchmarks show state-of-the-art accuracy, with an average gain of 1.4% over DVLA-RL. The project page is available at https://peacelwh.github.io/TPAMI27-DVLA-RLpp/.
Figures & tables
Fig. 1: Illustration of support-evidence contamination. (a) Intrinsic bird cues and incidental reed patterns coexist in a support image. (b) Comparing intrinsic and nuisance agreement keeps the feathers and rejects the reeds. The reed plume has positive relative evidence but falls below the rejection margin of 0.15, illustrating the need for explicit rejection.
Fig. 2: Overview of the proposed DVLA-RL++ framework. (A) A description generator produces intrinsic and nuisance semantic banks from the labeled supports of each class, and a frozen text encoder embeds them. (B) The intrinsic semantics guide local and global alignment inside the visual backbone through RL-gated attention. (C) Support tokens compete against the two banks, the retained evidence is mixed with an intrinsic anchor, and class-wise aggregation yields the prototypes used for nearest-prototype recognition. (D) During training only, a sampled gate policy and a frozen reference gate follow separate state trajectories on the same episode, and their paired utilities update the gate policy. Query labels enter only the training objectives and the evaluation metrics.
Fig. 3: Illustration of the RL-gated attention block. A support-derived gate combines image-guided and text-guided semantic responses. The fused context enriches image tokens through joint transformer processing, after which only image tokens are propagated.
Fig. 4: Illustration of complementary semantic purification. Competing intrinsic and nuisance descriptions guide sparse support-token allocation. An intrinsic semantic anchor fills the unallocated mass before class-wise aggregation and normalization.
Fig. 5: Illustration of counterfactual reinforcement-learning gating. Sampled and frozen reference policies propagate their own states on the same episode. Their paired utilities guide gate-policy updates while the representation remains fixed.
Method
Venue
miniImageNet
tieredImageNet
CIFAR-FS
1-shot
5-shot
1-shot
5-shot
1-shot
5-shot
ProtoNet [ 2 ]
NeurIPS 2017
62.39±0.21
80.53±0.14
68.23±0.23
84.03±0.16
72.20±0.70
83.50±0.50
AM3 [ 6 ]
NeurIPS 2019
65.30±0.49
78.10±0.36
69.08±0.47
82.58±0.31
-
-
DeepEMD [ 3 ]
CVPR 2020
65.91±0.82
82.41±0.56
71.16±0.87
86.03±0.58
-
-
SUN [ 31 ]
ECCV 2022
67.80±0.45
83.25±0.30
72.99±0.50
86.74±0.33
-
-
FewTURE [ 37 ]
NeurIPS 2022
68.02±0.88
84.51±0.53
72.96±0.92
86.43±0.67
72.80±0.88
86.14±0.64
TABLE I: Results (%) on miniImageNet, tieredImageNet, and CIFAR-FS under the 5-way 1-shot and 5-shot settings. Mean accuracy and 95% confidence intervals are reported. Bold indicates the highest mean accuracy.
Method
Venue
CUB
Stanford Cars
Stanford Dogs
1-shot
5-shot
1-shot
5-shot
1-shot
5-shot
ProtoNet [ 2 ]
NeurIPS 2017
63.44±0.56
83.17±0.35
45.01±0.49
87.19±0.31
41.61±0.50
76.78±0.36
FRN [ 49 ]
CVPR 2021
83.55±0.19
92.92±0.10
58.90±0.22
79.65±0.15
49.37±0.20
67.13±0.17
DAN [ 50 ]
AAAI 2022
72.89±0.50
86.60±0.31
70.21±0.50
85.55±0.31
59.81±0.50
77.19±0.35
MFGN [ 51 ]
IJCAI 2022
84.01±0.39
91.85±0.21
-
-
74.81±0.44
86.52±0.26
TDM [ 52 ]
CVPR 2022
84.36±0.19
93.37±0.10
68.36±0.22
86.14±0.13
57.64±0.22
75.03±0.16
TABLE II: Results (%) on CUB, Stanford Cars, and Stanford Dogs under the 5-way 1-shot and 5-shot settings. Mean accuracy and 95% confidence intervals are reported. Bold indicates the highest mean accuracy.
Method
Venue
CUB
Places
ChestX
1-shot
5-shot
1-shot
5-shot
1-shot
5-shot
GNN [ 56 ]
ICLR 2018
44.40±0.68
62.87±0.65
52.42±0.80
70.91±0.65
22.00±0.46
25.27±0.59
FWT [ 57 ]
ICLR 2020
45.50±0.46
64.97±0.68
53.44±0.79
70.70±0.67
22.04±0.44
25.18±0.45
AFA [ 58 ]
ECCV 2022
46.86±0.70
68.25±0.65
54.04±0.75
76.21±0.60
22.92±0.20
25.02±0.20
UCD [ 59 ]
NeurIPS 2022
40.65±0.68
58.54±0.70
51.84±0.72
72.19±0.60
22.64±0.40
26.26±0.45
ATA [ 60 ]
AIJ 2023
45.00±0.50
66.22±0.50
53.57±0.50
75.48±0.40
22.10±0.20
24.32±0.40
TABLE III: Cross-domain results (%) with training on miniImageNet and evaluation on CUB, Places, and ChestX. Mean accuracy and 95% confidence intervals are reported. Bold indicates the highest mean accuracy.
Fig. 6: Parameter sensitivity on miniImageNet (blue solid) and CUB (orange dashed) under the 5-way 1-shot setting. Panels vary (a) τs , (b) λ , and (c) γ . Accuracy changes are relative to each dataset’s default setting.
Variant
CSP
CRG
miniImageNet
CUB
Clean
Shift
Clean
Shift
DVLA-RL
×
×
81.7
75.8
91.9
85.0
+ CSP
✓
×
82.6
79.4
93.4
89.4
+ CRG
×
✓
82.4
78.4
93.0
88.1
DVLA-RL++
✓
✓
83.4
81.2
94.3
92.0
TABLE IV: Component ablations (%) on miniImageNet and CUB under the 5-way 1-shot setting. Clean and Shift denote accuracy before and after the matched context intervention.
Fig. 7: Signed-margin distributions of intrinsic (blue) and nuisance (orange) evidence on miniImageNet and CUB. Rows compare DVLA-RL and DVLA-RL++. OVL denotes histogram overlap; the dotted line marks the rejection boundary d−ξc=0 .
Variant
miniImageNet
CUB
Clean
Shift
Clean
Shift
DVLA-RL
81.7
75.8
91.9
85.0
Uniform allocation
82.2
77.6
92.8
87.2
Positive semantics only
82.7
79.0
93.5
89.1
No neutral option
83.0
79.7
93.8
90.0
Static gate
82.8
79.8
93.6
90.0
TABLE V: Internal design controls (%) on miniImageNet and CUB under the 5-way 1-shot setting. Clean and Shift follow the same protocol as Table IV .
Fig. 9: Attention maps of DVLA-RL, CSP alone, CRG alone, and DVLA-RL++ on nine miniImageNet images. Columns (a)–(i) cover object co-occurrence, occlusion, clutter, and camouflage.
Fig. 8: Policy diagnostics on miniImageNet. (a) Gradient-covariance trace normalized by the state-baseline estimator. (b) Frequencies of nuisance-cost increase and active PPO clipping for the state baseline (blue) and paired estimator (green).
Fig. 10: T-SNE visualization on novel classes from four datasets.
Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that exploiting few-shot image embeddings from a training set is effective for CLIP-based classification. In this work, we analyze mixing image and text prototypes from a bias-variance perspective and show that mixing prototypes acts like a variance shrinkage estimator. Naively mixing text and image prototypes combines two partially aligned spaces since the two modalities are not perfectly aligned. To address this, we project image prototypes onto the principal directions of the semantic text embedding space to obtain a task-semantic image subspace. Mixing the image prototypes with text embeddings in the task-semantic subspace improves few-shot classification. However, when the task-semantic subspace captures insufficient discriminative visual information, relying on this subspace alone can be suboptimal. On extensive experiments over several few-shot classification benchmarks, we show that combining a task-semantic mixed prototype classifier and an anisotropic image-specific classifier systematically outperforms existing methods.
Dipam Goswami, Simone Magistri, Gido M. van de Ven +4
Department of Computer Science, Universitat Autònoma de Barcelona, Spain · Computer Vision Center, Barcelona, Spain · Media Integration and Communication Center, University of Florence, Italy +3
Training-free few-shot adaptation methods have gained significant attention recently in the context of Vision-language Models (VLMs). Yet, current benchmarks rely on strong assumptions about the statistics of the adaptation data, e.g., class balance. We question these simplifying assumptions and introduce a more realistic benchmark that varies both the levels of class balance and the effective number of classes in few-shot tasks via Dirichlet sampling. Surprisingly, under our setting, we observe substantial drops in the performances of state-of-the-art methods, more so when the number of labeled samples increases. To mitigate this, we introduce PRiSM, a class-prototype regularization that can be deployed as a plug and play module on top of any existing baseline method, significantly improving performances. Our method optimizes a novel multi-term loss, which includes a regularizer maximizing inter-class pairwise distances, along with additional terms promoting support-feature alignment and fidelity to the baseline prototypes. Furthermore, we introduce an effective and computationally efficient block Majorize-Minimize optimizer for our objective. More specifically, we derive a valid blockwise Lipschitz constant (i.e., a bound on the Hessian's spectral norm), which can be computed efficiently via the Gershgorin circle theorem. Extensive experiments show that PRiSM improves several training-free baselines, with large gains when dealing with severe class imbalance and high numbers of classes.
Ghassen Baklouti, Omprakash Chakraborty, Jose Dolz +1
Few-shot object detection aims to detect novel object categories from only a few labeled examples, avoiding costly large-scale annotation. Recent prototype-based similarity learning approaches enable training-free adaptation by matching query features with class prototypes. However, they suffer from two fundamental limitations: (i) class confusion arising from inter-class similarity margin collapse, and (ii) insufficient visual cues for precise localization, as similarity scores capture only class-level semantic affinity while providing limited spatial information. To address these issues, we introduce two complementary components. Text-Anchored Semantic Mask (TSMa) leverages class-level text features as semantic anchors to identify semantically aligned channels through channel-wise interaction between visual and text features. By suppressing style-induced spurious responses and emphasizing class-intrinsic signals, TSMa enlarges inter-class similarity margins and mitigates class confusion. We further propose Stage-Aligned Hierarchical Autoregressive Regression (SHARe), which reformulates localization as a hierarchical autoregressive process that progressively refines bounding boxes across multiple stages. SHARe leverages the layer-wise characteristics of ViT representations by aligning feature abstraction levels with regression stages: deeper layers guide early coarse localization, while shallower layers rich in edge and texture cues refine spatial details in later stages. Experiments on COCO demonstrate a new state of the art, outperforming the previous best by +10.1 nAP, with extensive analysis validating each component. The code is available at https://github.com/VisualScienceLab-KHU/ReSet.