Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reasoning risking the degradation of pre-trained reasoning capabilities. We propose TTRSD, a test-time reinforcement learning framework combining multi-view answer-level self-distillation with visual contrastive token selection. A shared policy aggregates teacher predictions across original, cropped, and downsampled views into an answer distribution. Student trajectories generated from the original image receive rewards based on the support for their final answers in this distribution. To allocate this feedback precisely toward perceptual bottlenecks, we compare the log-probabilities of the same sampled tokens under original and visually ablated inputs while holding their textual prefixes fixed, selecting visually sensitive positions for policy-gradient updates. TTRSD separates update direction, determined by group-relative advantages, from update position, determined by visual sensitivity, without requiring ground-truth labels, external verifiers, or a separate teacher. With only 20 unlabeled adaptation samples, TTRSD improves performance across seven benchmarks and three VLMs, raising InternVL3-2B's MMMU accuracy from 35.79% to 49.32%(+13.53%), demonstrating cross-dataset generalization while preserving inherent reasoning integrity.
Figures & tables
Figure 1: Overview and performance of TTRSD . (a) TTRL derives rewards from majority-vote pseudo-labels and applies sequence-level advantages to all response tokens. (b) TTRSD improves average accuracy over TTRL across seven benchmarks on three VLM backbones. (c) TTRSD uses a shared policy to construct rewards from multi-view teacher answers and selects visually sensitive tokens through original–ablated prediction comparisons, enabling targeted policy-gradient updates.
Figure 2: Overview of TTRSD . (1) Multi-view answer-level self-distillation aggregates teacher predictions into an answer distribution to reward original-view student trajectories. (2) Visual contrastive token selection compares the log-probabilities of the same student tokens under original and ablated images, selecting visually sensitive positions. (3) Policy optimization combines group-relative advantages with token masks for GRPO updates, while applying KL regularization over all valid response tokens. Teacher and student share the policy parameters throughout adaptation.
Backbone
Method
WeMath
LogicVista
MathVista
MathVerse
MathVision
MMMU
MME-R
Avg.
Closed-source Model
GPT-4o
–
50.60
64.40
71.60
49.90
43.80
70.70
30.20
54.46
Gemini-2.0-Flash
–
47.42
–
70.46
43.65
47.82
69.30
–
–
Open-source Model
MiniCPM-V 2.6
–
38.62
27.54
60.60
38.30
23.40
49.80
19.34
–
InternLM- XComposer2-VL-7B
–
12.67
–
57.60
25.90
14.54
43.00
–
–
Table 1: Main results on seven vision–language benchmarks. We report accuracy (%) and the seven-benchmark average (Avg.). The best result within each adaptation backbone is shown in bold . Missing published results are denoted by “–”.
Figure 3: Cross-dataset generalization of TTRSD with InternVL3-8B. Bars compare zero-shot and source-adapted accuracy without target adaptation data. Arrows indicate transfer direction; gains are in percentage points.
MVSD
VCTS
MathVista
MathVision
LogicVista
✓
✓
66.94
27.87
39.25
–
✓
62.44
25.19
37.92
✓
–
61.78
27.02
31.15
–
–
61.43
23.76
23.48
Table 2: Ablation studies on InternVL3-2B. (a) Component ablation: MVSD denotes multi-view answer-level self-distillation, and VCTS denotes visual contrastive token selection. (b) Effect of the token selection ratio ρ on MathVista. All results are accuracy (%).
Figure 4: Multi-view teacher voting improves pseudo-label accuracy over individual views on both benchmarks.
Figure 5: Visual token selection in a table-reading example. Purple tokens are selected for policy-gradient updates based on their sensitivity to visual ablation. More visualizations in Appendix.
Benchmark
All
Selected
Unselected
WeMath
0.1709
0.7012
0.0367
LogicVista
0.4283
1.5862
0.1262
MathVista
0.3825
1.4089
0.0974
MathVerse
0.3538
1.1875
0.1286
MathVision
0.2036
0.7295
0.0677
MMMU
0.2839
1.0563
0.0837
Table 3: Token-level visual sensitivity at ρ=0.2 . Values are mean absolute log-probability differences between original and visually ablated inputs.
Benchmark
All
Selected
Unselected
WeMath
0.1709
0.7012
0.0367
LogicVista
0.4283
1.5862
0.1262
MathVista
0.3825
1.4089
0.0974
MathVerse
0.3538
1.1875
0.1286
MathVision
0.2036
0.7295
0.0677
MMMU
0.2839
1.0563
0.0837
Table 3: Token-level visual sensitivity at ρ=0.2 . Values are mean absolute log-probability differences between original and visually ablated inputs.
Selection rule
Acc. (%)
Δ (pp)
Visual (ours)
66.94
+5.16
Random
62.45
+0.67
High entropy
61.43
-0.35
All tokens
61.78
0.00
Table 4: Token-selection ablation on MathVista. Selection methods retain 20% of valid response tokens; other training settings are fixed.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Value
Actor learning rate
5×10−7
Training epochs
8
PPO epochs per batch
1
PPO clipping range
0.2
KL-loss coefficient
1×10−3
Training batch size
32
Appendix
Table 5: Training hyperparameters used in all experiments.
Benchmark
Accuracy (%)
LogicVista
57.58±0.17
MathVerse
67.01±0.16
MathVision
39.64±0.37
MathVista
80.93±0.75
WeMath
72.96±1.26
Appendix
Table 6: Multi-seed results of TTRSD with Qwen3-VL-4B. Accuracy (%) is reported as the mean and standard deviation over five independent random seeds.
Benchmark
20 Questions
100 Questions
Δ
WeMath
37.21
38.14
0.93
MMMU
49.32
50.21
0.89
LogicVista
39.25
41.78
2.53
Appendix
Table 7: Effect of scaling the unlabeled adaptation set from 20 to 100 questions with InternVL3-2B. Accuracy is reported in percent.
Benchmark
Zero-shot
Ours
Δ
DTD
37.12
87.94
+50.8
MMStar
47.97
51.11
+3.1
SEED-Bench
69.85
70.82
+1.0
RealWorldQA
63.75
64.57
+0.8
Appendix
Table 8: Results on perception-oriented and multi-disciplinary benchmarks with InternVL3-2B. Accuracy (%) is reported before and after adaptation.
Stage
Time
Share
Rollout generation (111 tokens, autoregressive)
13,170 ms
96.5%
Real-image scoring forward (no grad)
94.1 ms
0.7%
Blank-image scoring forward (no grad, ours)
94.2 ms
0.7%
Policy forward + backward
224.4 ms
1.6%
Optimizer update
63.4 ms
0.5%
Total
13,646 ms
100%
Appendix
Table 9: Wall-clock breakdown of one training step for a single sequence (InternVL3-2B, bfloat16, one NVIDIA A100 GPU; averaged over ten runs). The blank-image scoring pass is the only component added by our method.
Figure 6: Case Study 1: numeric extraction from a chart.
Figure 7: Case Study 2: object counting.
Figure 8: Case Study 3 (failure): a perceptual illusion.
School of Mathematics, Tianjin University, Tianjin, China · Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China · Shanghai Advanced Institute of Finance (SAIFS), East China Normal University, Shanghai, China +4