Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning
Organizations: Google DeepMind · UC San Diego · Work done at Google DeepMind
Abstract
Latent reasoning has advanced multimodal reasoning through a two-stage training paradigm: (1) a helper image is encoded into latent tokens to teach visual chain-of-thought during a supervised fine-tuning (SFT) stage, and (2) these latent tokens are further refined with reward feedback during a reinforcement learning (RL) stage. In this paper, we identify two key limitations of this framework, one in each stage. First, the SFT stage typically relies on an off-the-shelf vision encoder to encode the helper image, yielding suboptimal latent representations that may not be well aligned with the downstream reasoning task. Second, existing RL methods treat the latent component only through deterministic regularization, which constrains policy drift but does not create alternative latent trajectories for exploration. To address these limitations, we propose Scaffolding Minds. Our approach learns a dedicated scaffolding encoder that provides an optimized target in latent space, and learns both the mean and variance of the RL sampler. We further show that these two improvements are complementary, together yielding substantial gains over strong baselines. Empirically, our method improves over the strongest latent reasoning baseline by +9.5 points on FrozenLake spatial planning, with the gain widening to +19 points on the 32x32 grids, and by +5.6 points on average across nine visual-centric reasoning benchmarks.
Figures & tables
| Method | Lv. 8 ( ) | Lv. 16 ( ) | Lv. 24 ( ) | Lv. 32 ( ) | Avg. |
|---|---|---|---|---|---|
| Base model & supervised fine-tuning | |||||
| SFT | 84.0 | 65.0 | 41.0 | 18.0 | 52.0 |
| SFT+GRPO | 85.0 | 67.0 | 43.0 | 20.0 | 53.8 |
| Image-generation methods | |||||
| VPRL ⋄ ( Xu et al., 2025 ) | 88.0 | 62.0 | 48.0 | 35.0 | 58.3 |
| DiffThinker ⋄ ( He et al., 2025 ) | 92.0 | 68.0 | 53.0 | 44.0 | 64.3 |
| Method | V ⋆ | BLINK | MMVP | MMStar | CVB. | HRB-4K | HRB-8K | MME | Jig. | Avg. |
|---|---|---|---|---|---|---|---|---|---|---|
| Closed-source models | ||||||||||
| GPT-4o | 66.0 | 60.0 | 70.7 | 61.6 | 80.1 | 59.0 | 55.5 | 52.0 | 53.0 | 62.0 |
| GPT-4V | 58.0 | 58.3 | 51.0 | 56.0 | 69.5 | 56.5 | 52.0 | 45.0 | 47.0 | 54.8 |
| GPT-4o-mini | 57.0 | 53.6 | 56.0 | 54.8 | 68.0 | 56.0 | 51.0 | 37.4 | 46.0 | 53.3 |
| Claude 3.7-Sonnet | 72.0 | 56.6 | 64.0 | 65.1 | 80.0 | 68.0 | 64.0 | 49.0 | 52.0 | 63.4 |
| Base model & supervised fine-tuning | ||||||||||
| Latent sampling | Stage 1 target | |||
| Stage 2 recipe | Mean | Variance | Scaffolding enc. | Frozen target |
| Stage 1 checkpoint (no RL) | – | 72.0 | 57.8 | |
| Text-only GRPO | no latent sampling | 72.6 ( ) | 58.4 ( ) | |
| VLPO ( Wang et al., 2026 ) | latent regularization | 72.9 ( ) | 59.0 ( ) | |
| Learned mean only | learned | fixed | 73.6 ( ) | 60.4 ( ) |
| Learned variance only | zero | learned | 73.2 ( ) | 59.8 ( ) |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| L8 | L16 | L24 | L32 | Avg. | |
|---|---|---|---|---|---|
| 1 | 85.0 | 70.0 | 54.0 | 45.0 | 63.5 |
| 2 | 88.0 | 74.0 | 57.0 | 48.0 | 66.8 |
| 4 | 94.0 | 79.0 | 63.0 | 52.0 | 72.0 |
| 8 | 93.0 | 78.0 | 61.0 | 51.0 | 70.8 |
| 16 | 91.0 | 76.0 | 59.0 | 50.0 | 69.0 |
| L8 | L16 | L24 | L32 | Avg. | |
|---|---|---|---|---|---|
| 1 | 85.0 | 70.0 | 54.0 | 45.0 | 63.5 |
| 2 | 88.0 | 74.0 | 57.0 | 48.0 | 66.8 |
| 4 | 94.0 | 79.0 | 63.0 | 52.0 | 72.0 |
| 8 | 93.0 | 78.0 | 61.0 | 51.0 | 70.8 |
| 16 | 91.0 | 76.0 | 59.0 | 50.0 | 69.0 |
| Helper image | Scaffold. enc. | L8 | L16 | L24 | L32 | Avg. |
|---|---|---|---|---|---|---|
| Red-arrow overlay | 85.0 | 67.0 | 52.0 | 40.0 | 61.0 | |
| 92.0 | 76.0 | 58.0 | 46.0 | 68.0 | ||
| Value-function heatmap | 88.0 | 71.0 | 47.0 | 25.0 | 57.8 | |
| 94.0 | 79.0 | 63.0 | 52.0 | 72.0 |
| Variant | V ⋆ | BLINK | MMVP | MMStar | CVB. | HRB-4K | HRB-8K | MME | Jig. | Avg. | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Stage 1 | Frozen target | 82.4 | 57.5 | 71.3 | 67.8 | 79.0 | 72.0 | 68.3 | 45.2 | 52.8 | 66.3 |
| Scaffolding encoder | 87.4 | 64.8 | 73.3 | 72.1 | 86.4 | 73.6 | 71.1 | 53.4 | 56.0 | 70.9 | |
| Stage 2 | VLPO | 89.3 | 66.2 | 75.3 | 72.9 | 87.6 | 75.9 | 72.9 | 55.7 | 58.2 | 72.7 |
| Scaffolding RL | 90.6 | 67.2 | 76.7 | 73.5 | 88.4 | 77.5 | 74.1 | 57.3 | 59.7 | 73.9 |
| Preceding latent context | V ⋆ | MMVP | CVB. | HRB-4K | Mean | Forward passes | Training time |
|---|---|---|---|---|---|---|---|
| Target (teacher forcing, ours) | 87.4 | 73.3 | 86.4 | 73.6 | 80.2 | ||
| Predicted (self-predicted) | 87.1 | 72.9 | 86.2 | 73.3 | 79.9 |
| Metric | Text-only GRPO | VLPO | Scaffolding RL |
|---|---|---|---|
| Reward s.d. within a rollout group | 0.15 | 0.32 | 0.42 |
| Pass@4 / @8 / @16 (%) | 74 / 76 / 80 | 75 / 79 / 82 | 79 / 82 / 85 |
| Latent Gaussian scale | N/A | 0.05 ∗ (fixed) | 0.04–0.15 (learned) |
| L8 | L16 | L24 | L32 | Avg. | ||
|---|---|---|---|---|---|---|
| 1.0 | 0.0 | 88.0 | 72.0 | 54.0 | 45.0 | 64.8 |
| 1.0 | 0.5 | 94.0 | 79.0 | 63.0 | 52.0 | 72.0 |
| 1.0 | 1.0 | 92.0 | 77.0 | 60.0 | 51.0 | 70.0 |
| 0.5 | 1.0 | 93.0 | 78.0 | 62.0 | 52.0 | 71.3 |
| Qwen2.5-VL-3B (full FT) | Qwen2.5-VL-7B (LoRA ) | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Method | Lv. 8 | Lv. 16 | Lv. 24 | Lv. 32 | Avg. | Lv. 8 | Lv. 16 | Lv. 24 | Lv. 32 | Avg. |
| SFT | 75.0 | 60.0 | 38.0 | 14.0 | 46.8 | 70.0 | 52.0 | 32.0 | 13.0 | 41.8 |
| SFT+GRPO | 77.0 | 62.0 | 40.0 | 16.0 | 48.8 | 71.0 | 54.0 | 34.0 | 14.0 | 43.3 |
| LVR ( Li et al., 2025b ) | 78.0 | 63.0 | 41.0 | 18.0 | 50.0 | 73.0 | 56.0 | 36.0 | 16.0 | 45.3 |
| Mirage ( Yang et al., 2025b ) | 79.0 | 65.0 | 43.0 | 20.0 | 51.8 | 74.0 | 58.0 | 38.0 | 18.0 | 47.0 |
| VaLR ( Jeon et al., 2026 ) | 83.0 | 71.0 | 51.0 | 28.0 | 58.3 | 80.0 | 65.0 | 44.0 | 24.0 | 53.3 |
| FrozenLake (single block) | Visual-centric (interleaved) | ||||
| Method | Latency (ms) | Overhead | Latency (ms) | Overhead | |
| Qwen2.5-VL-7B (base) | — | 260 | — | 740 | — |
| Latent reasoning methods | |||||
| Mirage ( Yang et al., 2025b ) | 4 | 266 | – | – | |
| LVR ( Li et al., 2025b ) | 4 | 267 | 794 | ||
| VaLR ( Jeon et al., 2026 ) | 4 | 268 | 802 | ||