Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reasoning risking the degradation of pre-trained reasoning capabilities. We propose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillation with visual contrastive token selection. A shared policy aggregates teacher predictions across original, cropped, and downsampled views into an answer distribution. Student trajectories generated from the original image receive rewards based on the support for their final answers in this distribution. To allocate this feedback precisely toward perceptual bottlenecks, we compare the log-probabilities of the same sampled tokens under original and visually ablated inputs while holding their textual prefixes fixed, selecting visually sensitive positions for policy-gradient updates. TTRSD separates update direction, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher. With only 20 unlabeled adaptation samples, TTRSD improves performance across seven benchmarks and three VLMs, raising InternVL3-2B's MMMU accuracy from 35.79% to 49.32%(+13.53%), demonstrating cross-dataset generalization while preserving inherent reasoning integrity.
Figures & tables
Figure 1: Overview and performance of TTRSD . (a) TTRL derives rewards from majority-vote pseudo-labels and applies sequence-level advantages to all response tokens. (b) TTRSD improves average accuracy over TTRL across seven benchmarks on three VLM backbones. (c) TTRSD uses a shared policy to construct rewards from multi-view teacher answers and selects visually sensitive tokens through original–ablated prediction comparisons, enabling targeted policy-gradient updates.
Figure 2: Overview of TTRSD . (1) Multi-view answer-level self-distillation aggregates teacher predictions into an answer distribution to reward original-view student trajectories. (2) Visual contrastive token selection compares the log-probabilities of the same student tokens under original and ablated images, selecting visually sensitive positions. (3) Policy optimization combines group-relative advantages with token masks for GRPO updates, while applying KL regularization over all valid response tokens. Teacher and student share the policy parameters throughout adaptation.
Backbone
Method
WeMath
LogicVista
MathVista
MathVerse
MathVision
MMMU
MME-R
Avg.
Closed-source Model
GPT-4o
–
50.60
64.40
71.60
49.90
43.80
70.70
30.20
54.46
Gemini-2.0-Flash
–
47.42
–
70.46
43.65
47.82
69.30
–
–
Open-source Model
MiniCPM-V 2.6
–
38.62
27.54
60.60
38.30
23.40
49.80
19.34
–
InternLM- XComposer2-VL-7B
–
12.67
–
57.60
25.90
14.54
43.00
–
–
Table 1: Main results on seven vision–language benchmarks. We report accuracy (%) and the seven-benchmark average (Avg.). The best result within each adaptation backbone is shown in bold . Missing published results are denoted by “–”.
Figure 3: Cross-dataset generalization of TTRSD with InternVL3-8B. Bars compare zero-shot and source-adapted accuracy without target adaptation data. Arrows indicate transfer direction; gains are in percentage points.
MVSD
VCTS
MathVista
MathVision
LogicVista
✓
✓
66.94
27.87
39.25
–
✓
62.44
25.19
37.92
✓
–
61.78
27.02
31.15
–
–
61.43
23.76
23.48
Table 2: Ablation studies on InternVL3-2B. (a) Component ablation: MVSD denotes multi-view answer-level self-distillation, and VCTS denotes visual contrastive token selection. (b) Effect of the token selection ratio ρ on MathVista. All results are accuracy (%).
Figure 4: Multi-view teacher voting improves pseudo-label accuracy over individual views on both benchmarks.
Figure 5: Visual token selection in a table-reading example. Purple tokens are selected for policy-gradient updates based on their sensitivity to visual ablation. More visualizations in Appendix.
Benchmark
All
Selected
Unselected
WeMath
0.1709
0.7012
0.0367
LogicVista
0.4283
1.5862
0.1262
MathVista
0.3825
1.4089
0.0974
MathVerse
0.3538
1.1875
0.1286
MathVision
0.2036
0.7295
0.0677
MMMU
0.2839
1.0563
0.0837
Table 3: Token-level visual sensitivity at ρ=0.2 . Values are mean absolute log-probability differences between original and visually ablated inputs.
Benchmark
All
Selected
Unselected
WeMath
0.1709
0.7012
0.0367
LogicVista
0.4283
1.5862
0.1262
MathVista
0.3825
1.4089
0.0974
MathVerse
0.3538
1.1875
0.1286
MathVision
0.2036
0.7295
0.0677
MMMU
0.2839
1.0563
0.0837
Table 3: Token-level visual sensitivity at ρ=0.2 . Values are mean absolute log-probability differences between original and visually ablated inputs.
Selection rule
Acc. (%)
Δ (pp)
Visual (ours)
66.94
+5.16
Random
62.45
+0.67
High entropy
61.43
-0.35
All tokens
61.78
0.00
Table 4: Token-selection ablation on MathVista. Selection methods retain 20% of valid response tokens; other training settings are fixed.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Actor learning rate
5×10−7
Training epochs
8
PPO epochs per batch
1
PPO clipping range
0.2
KL-loss coefficient
1×10−3
Training batch size
32
Appendix
Table 5: Training hyperparameters used in all experiments.
Benchmark
Accuracy (%)
LogicVista
57.58±0.17
MathVerse
67.01±0.16
MathVision
39.64±0.37
MathVista
80.93±0.75
WeMath
72.96±1.26
Appendix
Table 6: Multi-seed results of TTRSD with Qwen3-VL-4B. Accuracy (%) is reported as the mean and standard deviation over five independent random seeds.
Benchmark
20 Questions
100 Questions
Δ
WeMath
37.21
38.14
0.93
MMMU
49.32
50.21
0.89
LogicVista
39.25
41.78
2.53
Appendix
Table 7: Effect of scaling the unlabeled adaptation set from 20 to 100 questions with InternVL3-2B. Accuracy is reported in percent.
Benchmark
Zero-shot
Ours
Δ
DTD
37.12
87.94
+50.8
MMStar
47.97
51.11
+3.1
SEED-Bench
69.85
70.82
+1.0
RealWorldQA
63.75
64.57
+0.8
Appendix
Table 8: Results on perception-oriented and multi-disciplinary benchmarks with InternVL3-2B. Accuracy (%) is reported before and after adaptation.
Stage
Time
Share
Rollout generation (111 tokens, autoregressive)
13,170 ms
96.5%
Real-image scoring forward (no grad)
94.1 ms
0.7%
Blank-image scoring forward (no grad, ours)
94.2 ms
0.7%
Policy forward + backward
224.4 ms
1.6%
Optimizer update
63.4 ms
0.5%
Total
13,646 ms
100%
Appendix
Table 9: Wall-clock breakdown of one training step for a single sequence (InternVL3-2B, bfloat16, one NVIDIA A100 GPU; averaged over ten runs). The blank-image scoring pass is the only component added by our method.
Figure 6: Case Study 1: numeric extraction from a chart.
Figure 7: Case Study 2: object counting.
Figure 8: Case Study 3 (failure): a perceptual illusion.
Vision-Language Models (VLMs) frequently suffer from visual perception errors and hallucinations that compromise answer accuracy in complex reasoning tasks. Reinforcement Learning with Verifiable Rewards (RLVR) offers a promising solution by optimizing policies using answer correctness signals. Despite their effectiveness, prevailing RLVR methods face two critical limitations. First, much of the sampling budget is wasted on trajectories doomed to fail due to early visual description errors. Second, sparse rewards cannot distinguish whether failures stem from visual perception or reasoning stages. We introduce MIRL, a decoupled framework that addresses both limitations by leveraging mutual information (MI) between generated descriptions and visual inputs as a cheap pre-screening signal. This enables intelligent budget allocation toward high-potential trajectories via forking, while decoupled training provides independent MI-based rewards for visual perception optimization, resolving reward blindness. Experiments on six vision-language reasoning benchmarks demonstrate that MIRL achieves 70.22% average accuracy and successfully surpasses the performance of sampling 16 complete trajectories using only 10 pre-samples with top-6 selection (25% fewer complete trajectories). Our code is available at: https://anonymous.4open.science/r/mirl-main/.
Yin Zhang, Jiaxuan Zhao, Zonghan Wu +5
School of Mathematics, Tianjin University, Tianjin, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China +4
Post-training enables vision-language models (VLMs) to understand human instructions and perform various downstream tasks. Current post-training methods usually rely on human-annotated data, distillation from external models, reinforcement learning with human feedback, or verifiable answers. This limits their ability to improve without external supervision. To tackle this, we propose NOPD (Noisy Student On-Policy Self-Distillation), a simple yet effective self-distillation approach that improves VLMs without any external models or ground-truth answers. Our key insight is that prediction discrepancies between clean and corrupted inputs naturally induce a self-supervision signal. In NOPD, the model learns from corrupted inputs while using its own predictions under clean inputs as token-level supervision. We show the effectiveness of NOPD on five visual reasoning tasks; it can match and even outperform reinforcement learning approaches or distillation from external models. Notably, when trained with 2.1K samples from Geometry3K, NOPD improves Qwen2.5-VL-7B by 20 points on its validation set. It also shows generalization on out-of-distribution test sets and achieves 7.4 point gains on MathVista. Furthermore, we demonstrate that NOPD is a general approach to enhance VLMs, achieving improvements across three models on 12 benchmarks.
Shuai Wang, Daoan Zhang, Zhe Tang +2
The Hong Kong University of Science and Technology (Guangzhou) · University of Rochester · Zhejiang University of Technology +1
Reinforcement learning (RL) has emerged as an effective paradigm for improving the reasoning capability of vision-language models (VLMs). However, RL-based optimization typically depends on costly high-quality annotations that are difficult to scale. Existing unsupervised alternatives may drift toward biased solutions due to weak visual grounding and the lack of reliable verification signals. We propose a self-evolving post-training framework, DUEL, where supervision emerges from adversarial interactions between two policies initialized from the same pretrained VLM. A Challenger generates an image-grounded true claim together with a minimally perturbed hard-negative counterpart, while a Solver verifies both claims against the image, encouraging fine-grained visual discrimination under near-neighbor semantics. To stabilize optimization, we introduce a length-normalized log-likelihood reward that preserves informative optimization signals beyond binary outcome supervision and improves learning stability under sparse feedback. Experiments show that DUEL consistently improves visual reasoning and robust discrimination without additional human annotations, external reward models, or image editing tools.