Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.
Figures & tables
Figure 1: Two Sources of Candidate Diversity. Temperature sampling (left) draws several outputs from the same standard computation path, starting from the initial multimodal representation h(0) . Architectural sampling (right) changes the computation path from the same h(0) before producing one candidate from each path. The background is a conceptual illustration of the answer space.
Figure 2: Architectural sampling scales more favorably with candidate budget. pass@k of Qwen3-VL-8B across six multimodal benchmarks as the candidate budget grows. Architectural sampling continues to improve candidate coverage as the pool grows, while temperature sampling plateaus earlier.
Figure 3: Architectural Sampling from a Frozen VLM. (left) The image and prompt form the initial multimodal context, to which generated tokens are appended. (right) The standard path applies every decoder layer once. Each architectural path repeats one contiguous window {\color[rgb]{0.1406,0.3164,0.6523}r} times before continuing through the remaining frozen layers in their original order. All paths share frozen weights; changing the window and {\color[rgb]{0.1406,0.3164,0.6523}r} yields the candidate pool. Snowflakes mark frozen weights.
CV-Bench
RealWorldQA
CountBench
MathVista
MMStar
AI2D
GQA
A-OKVQA
ScienceQA
CountQA
MMMU
BLINK
Avg.
Qwen2.5-VL-7B
Temperature
92.46
86.67
97.75
72.60
79.47
90.75
70.15
93.36
92.71
86.36
78.28
87.52
85.67
Architectural
95.72
89.80
98.82
75.20
86.80
95.52
76.50
97.03
97.37
86.99
83.15
94.07
89.75
Δ
▲ 3.26
▲ 3.13
▲ 1.07
▲ 2.60
▲ 7.33
▲ 4.77
▲ 6.35
▲ 3.67
▲ 4.66
▲ 0.63
▲ 4.87
▲ 6.55
▲ 4.07
Qwen3-VL-2B
Temperature
85.22
72.29
94.62
43.70
53.33
74.85
64.60
85.68
75.66
82.30
35.33
63.43
69.25
Table 1: Candidate Coverage Across Checkpoints and Benchmarks. pass@9 for nine temperature samples from the standard path and nine architectural candidates, one from each path, across five Qwen checkpoints and twelve benchmarks. All candidates use T=0.6 . The Δ row reports architectural sampling minus standard-path temperature sampling. Avg. is the mean across benchmarks, and the higher score for each checkpoint–benchmark pair is bold.
Table 5
Figure 4: Temperature Robustness and Candidate Diversity. Results on Qwen3-VL-8B over the five benchmarks. (left) Architectural sampling maintains stable pass@9 across decoding temperatures, including greedy decoding ( T=0 ). (right) Architectural candidates have higher 1−Self-BLEU than temperature samples. Density curves and individual observations show the two distributions, with diamonds marking their means ( 0.51 vs. 0.22 ).
CV-Bench
MMStar
AI2D
A-OKVQA
ScienceQA
Temperature
88.48
58.93
84.89
88.65
88.75
Beam Search
86.09
52.60
80.78
87.16
86.22
Top- k
89.12
60.27
85.65
89.61
89.74
Top- p
88.91
59.84
85.37
89.28
89.41
Architectural
96.40
82.73
93.48
95.72
97.97
Table 4: Comparison with Standard Decoding strategies. pass@9 for Qwen3-VL-8B using temperature sampling, beam search, top- k , top- p , and architectural sampling under equal candidate budgets. Architectural sampling is highlighted; the best score per benchmark is bold.
CV-Bench
MMStar
AI2D
A-OKVQA
ScienceQA
CountBench
Qwen3.5-VL-2B
Temperature rollouts
74.60
45.52
75.38
68.04
67.01
93.06
Architectural rollouts
80.29 ▲ 5.69
49.17 ▲ 3.66
76.94 ▲ 1.56
83.38 ▲ 15.34
77.33 ▲ 10.32
95.10 ▲ 2.04
Qwen3-VL-8B
Temperature rollouts
86.17
53.93
76.61
87.40
84.75
93.17
Architectural rollouts
86.29 ▲ 0.11
55.52 ▲ 1.59
77.80 ▲ 1.17
87.95 ▲ 0.55
88.82 ▲ 4.07
96.93 ▲ 3.76
Table 5: Architectural Rollouts for Test-Time Learning. pass@1 accuracy after TTRV on Qwen3.5-VL-2B and Qwen3-VL-8B. Architectural-rollout cells report the absolute change over temperature-sampled rollouts; the best setting per benchmark is shown in bold.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
CV-Bench
MMStar
AI2D
A-OKVQA
ScienceQA
Temperature
88.48±0.11
58.93±0.14
84.89±0.23
88.65±0.11
88.75±0.09
Architectural
96.40±0.18
82.73±0.12
93.48±0.31
95.72±0.19
97.97±0.25
Appendix
Table 7: Run-to-Run Variation. Mean ± standard deviation of pass@9 ; the best setting per benchmark is bold.
CV-Bench
MMStar
AI2D
A-OKVQA
ScienceQA
Avg.
InternVL3-2B
Temperature
94.62
85.07
93.55
94.50
96.28
92.80
Architectural
95.87
87.87
93.52
96.16
98.31
94.35
Δ
▲ 1.25
▲ 2.80
▼ 0.03
▲ 1.66
▲ 2.03
▲ 1.54
Gemma3-4B
Temperature
57.96
45.73
72.19
80.00
74.67
66.11
Appendix
Table 8: Generalization Beyond Qwen. pass@9 on InternVL3-2B and Gemma3-4B. Δ denotes architectural minus temperature sampling; the best score per benchmark is bold.
Figure 5: Candidate Coverage as the Pool Grows: Qwen2.5-VL-7B. pass@k across six benchmarks as the candidate budget grows for architectural and temperature sampling.
Figure 6: Candidate Coverage as the Pool Grows: Qwen3.5-VL-4B. pass@k across six benchmarks as the candidate budget grows for architectural and temperature sampling.
Window size ∣B∣
r=1
r=2
r=4
2
1.00×
1.06×
1.19×
3
1.00×
1.09×
1.28×
4
1.00×
1.13×
1.38×
5
1.00×
1.16×
1.47×
Appendix
Table 9: Relative Layer-Application Cost. For a decoder with L=32 , normalized to one standard decoder traversal. Cost grows linearly with window size ∣B∣ and recursion depth r ; the experimental setting ∣B∣=3 is highlighted.
Configuration
# cand.
Cost ( × )
Latency (s)
Δ vs. std. (s)
Standard pass (r=1)
1
1.00
5.12
+0.00
Window ∣B∣=3,r=2
4
1.09
5.60
+0.48
Window ∣B∣=3,r=4
4
1.28
6.56
+1.44
Arch. sampling (total)
9
10.5
53.8
−−
Temp. sampling (total)
9
9.0
46.1
−−
Appendix
Table 10: Wall-Clock Latency. Per-candidate and total latencies for Qwen3-VL-8B using architectural sampling versus an equal-sized temperature-sampling baseline. The full 9 -candidate architectural set is highlighted.
Test-time scaling improves the reasoning performance of large language models but incurs substantial cost in both total computation and latency. Existing adaptive sampling methods partially mitigate this issue by dynamically deciding when to stop sampling, yet they typically rely on heuristic rules or rely on distribution assumptions. In this work, we formulate adaptive sampling as a Markov decision process (MDP). We train a lightweight sampling controller with reinforcement learning (RL) to jointly balance answer correctness, latency, and computation cost. At each round, the controller decides to stop sampling or to acquire additional samples. Our method is lightweight which only relies on statistics of final answers, and can be trained and deployed on CPU. We further show that the resulting framework admits an interpretation as the Lagrangian relaxation of a constrained optimization problem with explicit budget constraints. Experiments against strong baselines such as ASC and ESC show that our method achieves improved trade-offs among answer correctness, sampling rounds, and total samples required.
Runpeng Dai, Tong Zheng, Rui Liu +2
University of North Carolina at Chapel Hill · University of Maryland, College Park · Washington University in St. Louis
Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that spend the same compute on easy and difficult prompts. These fixed budgets are also difficult to inspect because they do not explain why a given prompt receives a particular number of samples. We propose adaptive} test-time scaling with a lightweight fuzzy controller that maps interpretable signals, including estimated prompt complexity and model confidence, to a per-query sampling budget. The controller assigns fewer samples to easier or more confident prompts and more samples to harder or less certain prompts, making inference-time compute inspectable rather than fixed or opaque. We evaluate under a fair-alignment protocol with matched decoding settings and controlled answer selection, and compare against best-of-N, compute-aware scaling, and self-certainty-based baselines on question-answering and mathematical reasoning tasks. Across models and datasets, adaptive fuzzy control improves over several standard baselines and remains close to a selector-matched full-budget control while reducing the average number of samples. These findings suggest that interpretable adaptive sampling is a practical direction for more efficient test-time reasoning in large language models.
Test-time scaling with parallel branches is widely adopted to improve performance on challenging reasoning tasks. The predominant approach, repeated sampling, draws branches independently from a single policy, which can produce redundant attempts and thereby limit the gains from additional inference compute. To address this limitation, we propose Planned Test-Time Scaling (PTTS), which replaces independent sampling with a coordinated joint policy: a planner generates a solution outline for each branch, steering the branches toward distinct reasoning paths, and an executor produces a full solution conditioned on each outline. Formally, we show that PTTS strictly generalizes repeated sampling and, in a stylized setting, provably promotes coverage of complementary reasoning modes and yields better pass@k scaling. We instantiate PTTS on top of strong reasoning models, keeping them fixed as executors while replacing repeated sampling with PTTS inference to further enhance test-time scaling. Concretely, we develop two variants: PTTS-ZS prompts a model to jointly generate outlines for all branches in a single autoregressive pass, while PTTS-RL directly optimizes the planner against the pass@k reward using truncated execution rollouts for efficient training and a sharper reward signal. Across five mathematical reasoning benchmarks with Qwen3-1.7B and 4B, PTTS-ZS improves pass@64 over repeated sampling by up to 6.7 points, while PTTS-RL further increases the gain to up to 13.4 points. Further analysis indicates that broader coverage of distinct reasoning paths contributes to these gains. Overall, PTTS provides a general framework for improving test-time scaling by coordinating reasoning branches, with zero-shot and trainable instantiations that yield substantial performance gains.