Organizations: School of Software, Shandong University, Jinan 250101, China · Shenzhen Loop Area Institute, Shenzhen 518038, China · School of Computer and Artificial Intelligence, Shandong University of Finance and Economics, Jinan 250014, China · School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen), 518055, China
Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making support prototypes susceptible to contextual contamination. To address this problem, we propose DVLA-RL++, which extends DVLA-RL with complementary semantic purification (CSP) and counterfactual reinforcement-learning gating (CRG). Specifically, CSP generates intrinsic and nuisance descriptions from labeled supports and compares their agreement with each support token. An ambiguity-dependent rejection margin guides sparse evidence allocation, while an intrinsic semantic anchor fills the unassigned mass to provide a fallback when visual evidence is unreliable. CRG learns layer-wise semantic fusion strengths using a reward that balances recognition performance and nuisance exposure. An independently executed reference trajectory on the same episode provides a paired learning signal. Theoretical analysis relates retained evidence and anchor quality to prototype stability and establishes conditions for unbiased on-policy gradient estimation. Experiments on standard, fine-grained, and cross-domain benchmarks show state-of-the-art accuracy, with an average gain of 1.4% over DVLA-RL. The project page is available at https://peacelwh.github.io/TPAMI27-DVLA-RLpp/.
Figures & tables
Fig. 1: Illustration of support-evidence contamination. (a) Intrinsic bird cues and incidental reed patterns coexist in a support image. (b) Comparing intrinsic and nuisance agreement keeps the feathers and rejects the reeds. The reed plume has positive relative evidence but falls below the rejection margin of 0.15, illustrating the need for explicit rejection.
Fig. 2: Overview of the proposed DVLA-RL++ framework. (A) A description generator produces intrinsic and nuisance semantic banks from the labeled supports of each class, and a frozen text encoder embeds them. (B) The intrinsic semantics guide local and global alignment inside the visual backbone through RL-gated attention. (C) Support tokens compete against the two banks, the retained evidence is mixed with an intrinsic anchor, and class-wise aggregation yields the prototypes used for nearest-prototype recognition. (D) During training only, a sampled gate policy and a frozen reference gate follow separate state trajectories on the same episode, and their paired utilities update the gate policy. Query labels enter only the training objectives and the evaluation metrics.
Fig. 3: Illustration of the RL-gated attention block. A support-derived gate combines image-guided and text-guided semantic responses. The fused context enriches image tokens through joint transformer processing, after which only image tokens are propagated.
Fig. 4: Illustration of complementary semantic purification. Competing intrinsic and nuisance descriptions guide sparse support-token allocation. An intrinsic semantic anchor fills the unallocated mass before class-wise aggregation and normalization.
Fig. 5: Illustration of counterfactual reinforcement-learning gating. Sampled and frozen reference policies propagate their own states on the same episode. Their paired utilities guide gate-policy updates while the representation remains fixed.
Method
Venue
miniImageNet
tieredImageNet
CIFAR-FS
1-shot
5-shot
1-shot
5-shot
1-shot
5-shot
ProtoNet [ 2 ]
NeurIPS 2017
62.39±0.21
80.53±0.14
68.23±0.23
84.03±0.16
72.20±0.70
83.50±0.50
AM3 [ 6 ]
NeurIPS 2019
65.30±0.49
78.10±0.36
69.08±0.47
82.58±0.31
-
-
DeepEMD [ 3 ]
CVPR 2020
65.91±0.82
82.41±0.56
71.16±0.87
86.03±0.58
-
-
SUN [ 31 ]
ECCV 2022
67.80±0.45
83.25±0.30
72.99±0.50
86.74±0.33
-
-
FewTURE [ 37 ]
NeurIPS 2022
68.02±0.88
84.51±0.53
72.96±0.92
86.43±0.67
72.80±0.88
86.14±0.64
TABLE I: Results (%) on miniImageNet, tieredImageNet, and CIFAR-FS under the 5-way 1-shot and 5-shot settings. Mean accuracy and 95% confidence intervals are reported. Bold indicates the highest mean accuracy.
Method
Venue
CUB
Stanford Cars
Stanford Dogs
1-shot
5-shot
1-shot
5-shot
1-shot
5-shot
ProtoNet [ 2 ]
NeurIPS 2017
63.44±0.56
83.17±0.35
45.01±0.49
87.19±0.31
41.61±0.50
76.78±0.36
FRN [ 49 ]
CVPR 2021
83.55±0.19
92.92±0.10
58.90±0.22
79.65±0.15
49.37±0.20
67.13±0.17
DAN [ 50 ]
AAAI 2022
72.89±0.50
86.60±0.31
70.21±0.50
85.55±0.31
59.81±0.50
77.19±0.35
MFGN [ 51 ]
IJCAI 2022
84.01±0.39
91.85±0.21
-
-
74.81±0.44
86.52±0.26
TDM [ 52 ]
CVPR 2022
84.36±0.19
93.37±0.10
68.36±0.22
86.14±0.13
57.64±0.22
75.03±0.16
TABLE II: Results (%) on CUB, Stanford Cars, and Stanford Dogs under the 5-way 1-shot and 5-shot settings. Mean accuracy and 95% confidence intervals are reported. Bold indicates the highest mean accuracy.
Method
Venue
CUB
Places
ChestX
1-shot
5-shot
1-shot
5-shot
1-shot
5-shot
GNN [ 56 ]
ICLR 2018
44.40±0.68
62.87±0.65
52.42±0.80
70.91±0.65
22.00±0.46
25.27±0.59
FWT [ 57 ]
ICLR 2020
45.50±0.46
64.97±0.68
53.44±0.79
70.70±0.67
22.04±0.44
25.18±0.45
AFA [ 58 ]
ECCV 2022
46.86±0.70
68.25±0.65
54.04±0.75
76.21±0.60
22.92±0.20
25.02±0.20
UCD [ 59 ]
NeurIPS 2022
40.65±0.68
58.54±0.70
51.84±0.72
72.19±0.60
22.64±0.40
26.26±0.45
ATA [ 60 ]
AIJ 2023
45.00±0.50
66.22±0.50
53.57±0.50
75.48±0.40
22.10±0.20
24.32±0.40
TABLE III: Cross-domain results (%) with training on miniImageNet and evaluation on CUB, Places, and ChestX. Mean accuracy and 95% confidence intervals are reported. Bold indicates the highest mean accuracy.
Fig. 6: Parameter sensitivity on miniImageNet (blue solid) and CUB (orange dashed) under the 5-way 1-shot setting. Panels vary (a) τs , (b) λ , and (c) γ . Accuracy changes are relative to each dataset’s default setting.
Variant
CSP
CRG
miniImageNet
CUB
Clean
Shift
Clean
Shift
DVLA-RL
×
×
81.7
75.8
91.9
85.0
+ CSP
✓
×
82.6
79.4
93.4
89.4
+ CRG
×
✓
82.4
78.4
93.0
88.1
DVLA-RL++
✓
✓
83.4
81.2
94.3
92.0
TABLE IV: Component ablations (%) on miniImageNet and CUB under the 5-way 1-shot setting. Clean and Shift denote accuracy before and after the matched context intervention.
Fig. 7: Signed-margin distributions of intrinsic (blue) and nuisance (orange) evidence on miniImageNet and CUB. Rows compare DVLA-RL and DVLA-RL++. OVL denotes histogram overlap; the dotted line marks the rejection boundary d−ξc=0 .
Variant
miniImageNet
CUB
Clean
Shift
Clean
Shift
DVLA-RL
81.7
75.8
91.9
85.0
Uniform allocation
82.2
77.6
92.8
87.2
Positive semantics only
82.7
79.0
93.5
89.1
No neutral option
83.0
79.7
93.8
90.0
Static gate
82.8
79.8
93.6
90.0
TABLE V: Internal design controls (%) on miniImageNet and CUB under the 5-way 1-shot setting. Clean and Shift follow the same protocol as Table IV .
Fig. 9: Attention maps of DVLA-RL, CSP alone, CRG alone, and DVLA-RL++ on nine miniImageNet images. Columns (a)–(i) cover object co-occurrence, occlusion, clutter, and camouflage.
Fig. 8: Policy diagnostics on miniImageNet. (a) Gradient-covariance trace normalized by the state-baseline estimator. (b) Frequencies of nuisance-cost increase and active PPO clipping for the state baseline (blue) and paired estimator (green).
Fig. 10: T-SNE visualization on novel classes from four datasets.
Department of Computer Science, Universitat Autònoma de Barcelona, Spain · Computer Vision Center, Barcelona, Spain · Media Integration and Communication Center, University of Florence, Italy +3