Token-Disentangled Latent Test-Time Scaling for Vision-Language Reasoning
Organizations: Qwen Business Unit of Alibaba · School of Artificial Intelligence, Nanjing University · National Key Laboratory for Novel Software Technology, Nanjing University · Zhejiang University · University of Waterloo · Chinese Academy of Sciences · The Hong Kong University of Science and Technology · Frontier Robotics
Abstract
Latent test-time scaling improves reasoning by refining hidden states during inference, but existing methods typically apply a single scalar reward to all editable latent tokens. For multimodal large language models, this global update ignores that generated tokens play different roles: some are sensitive to visual evidence, while others correspond to uncertain reasoning decisions. We present Token-Disentangled Latent Test-Time Scaling, an inference-time framework that makes latent refinement token-role-aware. Starting from an initial generated trajectory, we optimize a short hidden-state prefix while routing perception-side visual feedback to image-sensitive tokens and reasoning feedback to high-entropy tokens. Tokens selected by neither route are constrained by an anchor regularizer. Across both perception and reasoning benchmarks on Qwen2.5-VL-7B and InternVL3.5-8B, our method lifts macro accuracy over CoT by +2.57 and +1.51 respectively, and outperforms strong output-space test-time scaling baselines under matched decoded-candidate budgets. Code is available at https://github.com/Qwen-Applications/TD-LTTS.
Figures & tables
| Method | Perception | Reasoning | ||||||
|---|---|---|---|---|---|---|---|---|
| MMStar | RWQA | Hallusion | Avg. | ScienceQA | MathVista | LogicVista | Avg. | |
| Qwen2.5-VL-7B | ||||||||
| CoT | 62.00 | 63.66 | 70.56 | 65.41 | 89.74 | 67.80 | 44.52 | 67.35 |
| Self-consistency | 63.47 | 65.49 | 68.77 | 65.91 | 89.60 | 70.90 | 42.51 | 67.67 |
| Best-of-N | 63.93 | 66.37 | 70.56 | 66.95 | 90.63 | 71.60 | 42.73 | 68.32 |
| Reward-only | 64.67 | 67.06 | 70.45 | 67.39 | 90.12 | 70.90 | 42.95 | 67.99 |
| Base Model | Perception | Reasoning | ||||||
|---|---|---|---|---|---|---|---|---|
| MMStar | RWQA | Hallusion | Avg. | ScienceQA | MathVista | LogicVista | Avg. | |
| Qwen2.5-VL-3B ( Bai et al., 2025 ) | 59.50 | 54.07 | 63.30 | 58.96 | 80.76 | 63.14 | 40.93 | 61.61 |
| Ours | 62.20 2.70 | 56.13 2.06 | 65.30 2.00 | 61.21 2.25 | 80.91 0.15 | 64.05 0.91 | 41.88 0.95 | 62.28 0.67 |
| InternVL3.5-4B ( Wang et al., 2025 ) | 69.10 | 64.87 | 63.62 | 65.86 | 93.75 | 54.77 | 41.36 | 63.29 |
| Ours | 70.70 1.60 | 65.13 0.26 | 64.04 0.42 | 66.62 0.76 | 94.86 1.11 | 63.53 8.76 | 43.04 1.68 | 67.14 3.85 |
| LLaVA-OV-1.5-8B ( An et al., 2025 ) | 62.67 | 66.80 | 61.72 | 63.73 | 90.68 | 68.30 | 43.85 | 67.61 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
| LogicVista | RealWorldQA | |
|---|---|---|
| 45.10 (+0.58) | 66.40 (+2.74) | |
| 45.55 (+1.03) | 66.80 (+3.14) | |
| 46.31 (+1.79) | 67.19 (+3.53) | |
| 46.98 (+2.46) | 67.45 (+3.79) | |
| 46.53 (+2.01) | 67.06 (+3.40) | |
| 45.86 (+1.34) | 66.80 (+3.14) |
| Setting | LogicVista | RealWorldQA |
|---|---|---|
| strict disjoint (default) | 46.98 (+2.46) | 67.45 (+3.79) |
| overlap allowed | 46.31 (+1.79) | 67.06 (+3.40) |
| Setting | LogicVista | RealWorldQA |
|---|---|---|
| correctly routed (default) | 46.98 (+2.46) | 67.45 (+3.79) |
| swapped routing | 45.41 (+0.89) | 66.27 (+2.61) |
| Field | Routed rollout excerpt |
|---|---|
| Question | How many dogs can be seen in the image? Options: A: , B: , C: , D: . |
| Initial rollout | I need to identify all the dogs in the image. There is one dog visible on the left side of the image, lying down. No other dogs are clearly visible in the rest of the room. Final answer: C. |
| Selected refined candidate | I need to carefully identify the dogs in the image. There is a dog visible on the left side of the image, lying down. Another dog is partially visible behind the couch near the center of the room. No other dogs are clearly visible. Final answer: B. |
| Routing signal | Image-sensitivity top- /mean ; entropy top- /mean . |
| Case | Query summary | Ground truth | Initial Final | Qualitative change |
|---|---|---|---|---|
| MMStar, object counting | How many dogs can be seen in the image? | B: | C B | The initial rollout counts one visible dog. The refined candidate adds a second, partially visible dog behind the couch. |
| RealWorldQA, scene geometry | What level is the ground at? Options: flat, incline, decline. | B: incline | A B | The initial rollout treats the street as flat. The refined answer uses the slope toward the horizon and predicts incline. |
| MathVista, chart reading | How many bars have values larger than ? | The initial rollout counts both bars. The refined answer keeps only the bar above the threshold and rejects the bar below . | ||
| LogicVista, mechanical reasoning | If the weight is lifted by mm, which pulley rope must be pulled further? | C | B C | The initial rollout selects the simpler two-pulley system. The refined answer identifies the system requiring the longer rope displacement. |
| HallusionBench, temporal order | The plug is removed from the power outlet. Are the images in the correct positive order? | No | Yes No | The initial rollout assumes an insertion sequence. The refined answer rejects the sequence as inconsistent with the stated removal event. |