Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computations by reusing selected blocks of decoder layers. Varying the block location and repetition count introduces computational diversity without updating model weights or adding auxiliary parameters. Across five Qwen checkpoints and twelve multimodal benchmarks, architectural sampling improves pass@9 over standard-path temperature sampling by 6.58 percentage points on average at the same nine-candidate budget. Reusing early layers yields the strongest gains, and the improvement in candidate coverage persists even under greedy decoding. The resulting candidates show lower lexical overlap and improve accuracy when used as rollouts for label-free test-time reinforcement learning. These findings extend the benefits of our architectural sampling beyond candidate coverage, demonstrating more effective learning from a model's own outputs.
Figures & tables
Figure 1: Two Sources of Candidate Diversity. Temperature sampling (left) draws several outputs from the same standard computation path, starting from the initial multimodal representation h(0) . Architectural sampling (right) changes the computation path from the same h(0) before producing one candidate from each path. The background is a conceptual illustration of the answer space.
Figure 2: Architectural sampling scales more favorably with candidate budget. pass@k of Qwen3-VL-8B across six multimodal benchmarks as the candidate budget grows. Architectural sampling continues to improve candidate coverage as the pool grows, while temperature sampling plateaus earlier.
Figure 3: Architectural Sampling from a Frozen VLM. (left) The image and prompt form the initial multimodal context, to which generated tokens are appended. (right) The standard path applies every decoder layer once. Each architectural path repeats one contiguous window {\color[rgb]{0.1406,0.3164,0.6523}r} times before continuing through the remaining frozen layers in their original order. All paths share frozen weights; changing the window and {\color[rgb]{0.1406,0.3164,0.6523}r} yields the candidate pool. Snowflakes mark frozen weights.
CV-Bench
RealWorldQA
CountBench
MathVista
MMStar
AI2D
GQA
A-OKVQA
ScienceQA
CountQA
MMMU
BLINK
Avg.
Qwen2.5-VL-7B
Temperature
92.46
86.67
97.75
72.60
79.47
90.75
70.15
93.36
92.71
86.36
78.28
87.52
85.67
Architectural
95.72
89.80
98.82
75.20
86.80
95.52
76.50
97.03
97.37
86.99
83.15
94.07
89.75
Δ
▲ 3.26
▲ 3.13
▲ 1.07
▲ 2.60
▲ 7.33
▲ 4.77
▲ 6.35
▲ 3.67
▲ 4.66
▲ 0.63
▲ 4.87
▲ 6.55
▲ 4.07
Qwen3-VL-2B
Temperature
85.22
72.29
94.62
43.70
53.33
74.85
64.60
85.68
75.66
82.30
35.33
63.43
69.25
Table 1: Candidate Coverage Across Checkpoints and Benchmarks. pass@9 for nine temperature samples from the standard path and nine architectural candidates, one from each path, across five Qwen checkpoints and twelve benchmarks. All candidates use T=0.6 . The Δ row reports architectural sampling minus standard-path temperature sampling. Avg. is the mean across benchmarks, and the higher score for each checkpoint–benchmark pair is bold.
Table 5
Figure 4: Temperature Robustness and Candidate Diversity. Results on Qwen3-VL-8B over the five benchmarks. (left) Architectural sampling maintains stable pass@9 across decoding temperatures, including greedy decoding ( T=0 ). (right) Architectural candidates have higher 1−Self-BLEU than temperature samples. Density curves and individual observations show the two distributions, with diamonds marking their means ( 0.51 vs. 0.22 ).
CV-Bench
MMStar
AI2D
A-OKVQA
ScienceQA
Temperature
88.48
58.93
84.89
88.65
88.75
Beam Search
86.09
52.60
80.78
87.16
86.22
Top- k
89.12
60.27
85.65
89.61
89.74
Top- p
88.91
59.84
85.37
89.28
89.41
Architectural
96.40
82.73
93.48
95.72
97.97
Table 4: Comparison with Standard Decoding strategies. pass@9 for Qwen3-VL-8B using temperature sampling, beam search, top- k , top- p , and architectural sampling under equal candidate budgets. Architectural sampling is highlighted; the best score per benchmark is bold.
CV-Bench
MMStar
AI2D
A-OKVQA
ScienceQA
CountBench
Qwen3.5-VL-2B
Temperature rollouts
74.60
45.52
75.38
68.04
67.01
93.06
Architectural rollouts
80.29 ▲ 5.69
49.17 ▲ 3.66
76.94 ▲ 1.56
83.38 ▲ 15.34
77.33 ▲ 10.32
95.10 ▲ 2.04
Qwen3-VL-8B
Temperature rollouts
86.17
53.93
76.61
87.40
84.75
93.17
Architectural rollouts
86.29 ▲ 0.11
55.52 ▲ 1.59
77.80 ▲ 1.17
87.95 ▲ 0.55
88.82 ▲ 4.07
96.93 ▲ 3.76
Table 5: Architectural Rollouts for Test-Time Learning. pass@1 accuracy after TTRV on Qwen3.5-VL-2B and Qwen3-VL-8B. Architectural-rollout cells report the absolute change over temperature-sampled rollouts; the best setting per benchmark is shown in bold.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
CV-Bench
MMStar
AI2D
A-OKVQA
ScienceQA
Temperature
88.48±0.11
58.93±0.14
84.89±0.23
88.65±0.11
88.75±0.09
Architectural
96.40±0.18
82.73±0.12
93.48±0.31
95.72±0.19
97.97±0.25
Appendix
Table 7: Run-to-Run Variation. Mean ± standard deviation of pass@9 ; the best setting per benchmark is bold.
CV-Bench
MMStar
AI2D
A-OKVQA
ScienceQA
Avg.
InternVL3-2B
Temperature
94.62
85.07
93.55
94.50
96.28
92.80
Architectural
95.87
87.87
93.52
96.16
98.31
94.35
Δ
▲ 1.25
▲ 2.80
▼ 0.03
▲ 1.66
▲ 2.03
▲ 1.54
Gemma3-4B
Temperature
57.96
45.73
72.19
80.00
74.67
66.11
Appendix
Table 8: Generalization Beyond Qwen. pass@9 on InternVL3-2B and Gemma3-4B. Δ denotes architectural minus temperature sampling; the best score per benchmark is bold.
Figure 5: Candidate Coverage as the Pool Grows: Qwen2.5-VL-7B. pass@k across six benchmarks as the candidate budget grows for architectural and temperature sampling.
Figure 6: Candidate Coverage as the Pool Grows: Qwen3.5-VL-4B. pass@k across six benchmarks as the candidate budget grows for architectural and temperature sampling.
Window size ∣B∣
r=1
r=2
r=4
2
1.00×
1.06×
1.19×
3
1.00×
1.09×
1.28×
4
1.00×
1.13×
1.38×
5
1.00×
1.16×
1.47×
Appendix
Table 9: Relative Layer-Application Cost. For a decoder with L=32 , normalized to one standard decoder traversal. Cost grows linearly with window size ∣B∣ and recursion depth r ; the experimental setting ∣B∣=3 is highlighted.
Configuration
# cand.
Cost ( × )
Latency (s)
Δ vs. std. (s)
Standard pass (r=1)
1
1.00
5.12
+0.00
Window ∣B∣=3,r=2
4
1.09
5.60
+0.48
Window ∣B∣=3,r=4
4
1.28
6.56
+1.44
Arch. sampling (total)
9
10.5
53.8
−−
Temp. sampling (total)
9
9.0
46.1
−−
Appendix
Table 10: Wall-Clock Latency. Per-candidate and total latencies for Qwen3-VL-8B using architectural sampling versus an equal-sized temperature-sampling baseline. The full 9 -candidate architectural set is highlighted.