Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning
Authors: Hao-Xuan Ma, Yihao Liu, Yutao Sun, Yanting Miao, Mengyu Zhou, YiCheng Xiao, Long Chen, Zhenguo Li, +3 more
Organizations: Qwen Business Unit of Alibaba · School of Artificial Intelligence, Nanjing University · National Key Laboratory for Novel Software Technology, Nanjing University · Zhejiang University · University of Waterloo · Chinese Academy of Sciences · The Hong Kong University of Science and Technology · Frontier Robotics
Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evidence, while others correspond to uncertain reasoning decisions. We present Token-Disentangled Latent Test-Time Scaling, an inference-time framework that makes latent refinement token-role-aware. Starting from an initial generated trajectory, we optimize a short hidden-state prefix while routing perception-side visual feedback to image-sensitive tokens and reasoning feedback to high-entropy tokens. Tokens selected by neither route are constrained by an anchor regularizer. Across both perception and reasoning benchmarks on Qwen2.5-VL-7B and InternVL3.5-8B, our method lifts macro accuracy over CoT by +2.57 and +1.51 respectively, and outperforms strong output-space test-time scaling baselines under matched decoded-candidate budgets. Code is available at https://github.com/Qwen-Applications/TD-LTTS.
Figures & tables
Figure 1: From global latent refinement to token-disentangled latent refinement. Our method augments latent test-time scaling with token-role-aware routing, separate reasoning and visual rewards, and anchor regularization, improving visual engagement, answer discrimination, and consistency across multimodal reasoning benchmarks.
Figure 2: Framework of Token-Disentangled Latent Test-Time Scaling. A frozen VLM first produces an initial chain-of-thought response and its hidden-state trajectory. We refine an early latent prefix from this trajectory at test time. Visual tokens (identified by high image sensitivity) receive visual rewards, while reasoning tokens (identified by high entropy) receive reasoning rewards. Token routing prevents visual and reasoning feedback from collapsing into a single global latent update.
Method
Perception
Reasoning
MMStar
RWQA
Hallusion
Avg.
ScienceQA
MathVista
LogicVista
Avg.
Qwen2.5-VL-7B
CoT
62.00
63.66
70.56
65.41
89.74
67.80
44.52
67.35
Self-consistency
63.47
65.49
68.77
65.91
89.60
70.90
42.51
67.67
Best-of-N
63.93
66.37
70.56
66.95
90.63
71.60
42.73
68.32
Reward-only
64.67
67.06
70.45
67.39
90.12
70.90
42.95
67.99
Table 1: Main results on multimodal reasoning benchmarks. Best results in each model block are highlighted in bold , and second-best results are underlined.
Base Model
Perception
Reasoning
MMStar
RWQA
Hallusion
Avg.
ScienceQA
MathVista
LogicVista
Avg.
Qwen2.5-VL-3B ( Bai et al., 2025 )
59.50
54.07
63.30
58.96
80.76
63.14
40.93
61.61
Ours
62.20 ↑ 2.70
56.13 ↑ 2.06
65.30 ↑ 2.00
61.21 ↑ 2.25
80.91 ↑ 0.15
64.05 ↑ 0.91
41.88 ↑ 0.95
62.28 ↑ 0.67
InternVL3.5-4B ( Wang et al., 2025 )
69.10
64.87
63.62
65.86
93.75
54.77
41.36
63.29
Ours
70.70 ↑ 1.60
65.13 ↑ 0.26
64.04 ↑ 0.42
66.62 ↑ 0.76
94.86 ↑ 1.11
63.53 ↑ 8.76
43.04 ↑ 1.68
67.14 ↑ 3.85
LLaVA-OV-1.5-8B ( An et al., 2025 )
62.67
66.80
61.72
63.73
90.68
68.30
43.85
67.61
Table 2: Performance comparison on perception and reasoning benchmarks. ↑ and ↓ indicate performance changes compared with the base model.
Figure 3: Component ablation. Perception is the average of MMStar and RealWorldQA; reasoning is the average of MathVista and LogicVista (two representative tasks per group, used as a cost-controlled subset of the Table 1 groups).
Figure 4: Effect of the maximum number of latent refinement steps K on the three-benchmark perception and reasoning averages from Table 1 .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
ρv=ρr
LogicVista
RealWorldQA
0.10
45.10 (+0.58)
66.40 (+2.74)
0.20
45.55 (+1.03)
66.80 (+3.14)
0.30
46.31 (+1.79)
67.19 (+3.53)
0.40
46.98 (+2.46)
67.45 (+3.79)
0.50
46.53 (+2.01)
67.06 (+3.40)
0.60
45.86 (+1.34)
66.80 (+3.14)
Appendix
Table 3: Diagonal top- k ratio sweep with ρv=ρr .
Setting
LogicVista
RealWorldQA
strict disjoint (default)
46.98 (+2.46)
67.45 (+3.79)
overlap allowed
46.31 (+1.79)
67.06 (+3.40)
Appendix
Table 4: Strict disjoint masks vs. overlapping masks at ρ=0.4 .
Setting
LogicVista
RealWorldQA
correctly routed (default)
46.98 (+2.46)
67.45 (+3.79)
swapped routing
45.41 (+0.89)
66.27 (+2.61)
Appendix
Table 6: Swapped-routing sanity check at ρ=0.4 .
Field
Routed rollout excerpt
Question
How many dogs can be seen in the image? Options: A: 3 , B: 2 , C: 1 , D: 4 .
Initial rollout
I need to identify all the dogs in the image. There is one dog visible on the left side of the image, lying down. No other dogs are clearly visible in the rest of the room. Final answer: C.
Selected refined candidate
I need to carefully identify the dogs in the image. There is a dog visible on the left side of the image, lying down. Another dog is partially visible behind the couch near the center of the room. No other dogs are clearly visible. Final answer: B.
Routing signal
Image-sensitivity top- k /mean =1.79/0.32 ; entropy top- k /mean =1.83/0.54 .
Appendix
Table 10: Readable span-level visualization of token routing for the MMStar dog-counting case. Color legend follows the surrounding paragraph.
Case
Query summary
Ground truth
Initial → Final
Qualitative change
MMStar, object counting
How many dogs can be seen in the image?
B: 2
C → B
The initial rollout counts one visible dog. The refined candidate adds a second, partially visible dog behind the couch.
RealWorldQA, scene geometry
What level is the ground at? Options: flat, incline, decline.
B: incline
A → B
The initial rollout treats the street as flat. The refined answer uses the slope toward the horizon and predicts incline.
MathVista, chart reading
How many bars have values larger than 100 ?
1
2→1
The initial rollout counts both bars. The refined answer keeps only the bar above the threshold and rejects the bar below 102 .
LogicVista, mechanical reasoning
If the weight is lifted by 10 mm, which pulley rope must be pulled further?
C
B → C
The initial rollout selects the simpler two-pulley system. The refined answer identifies the system requiring the longer rope displacement.
HallusionBench, temporal order
The plug is removed from the power outlet. Are the images in the correct positive order?
No
Yes → No
The initial rollout assumes an insertion sequence. The refined answer rejects the sequence as inconsistent with the stated removal event.
Appendix
Table 17: Successful case studies. Each row shows an example where the initial rollout is incorrect and the token-disentangled update recovers the correct short answer.