An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual question-answering benchmarks, we find that 89.4-98.1% of the rationale-side canvas remains unwritten at stabilization in single-block, EOS-suppressed LaViDa runs. On VBench, reducing block length from 128 to 8 changes this fraction from 89.4% to 1.7%, together with answer coverage and the eligible observation window. Under EOS-enabled prompting, direct instructions improve Nemotron's overall accuracy by 15.0 and 19.5 percentage points on M3CoT and ScienceQA, but reduce LaViDa/VBench accuracy by 11.0 points. A symmetric decomposition associates the larger absolute component of each change with coverage rather than conditional accuracy. Matched-canvas image ablations measure visual sensitivity alongside answer stabilization, separating the two temporal readouts. Together, these measurements distinguish answer stabilization, rationale unfolding, and visual sensitivity, and identify coverage as the larger component of the prompting differences.
Figures & tables
Figure 1: Overview of the measurement protocol. (a) Hold the model and input fixed, vary the block schedule, and record the decoding trace. (b) Distinguish stabilization of the logged answer candidate, Tstab , from permanent token commitment, Tcommit ; the image-ablation peak TPDM is a separate diagnostic. (c-i) Measure the unwritten fraction U of the rationale-side region R at Tstab . (c-ii) Compare image-present and image-removed token distributions on the same pre-step canvas and mask pattern. The PDM readout region is defined in Section 3.3 . The trace and curve are schematic illustrations. After commitment, a logged candidate may be copied from the canvas. Here PDMt=Dt in Equation 3 .
Data
N
rpos
Tstab
Tcommit
U (%)
C (%)
A (%)
V*Bench
191
0.15
6.41
23.65
89.4
100.0
41.9
M 3 CoT
400
0.09
0.93
7.60
98.1
99.8
71.5
ScienceQA
400
0.19
1.92
12.14
96.2
97.5
62.5
Table 1: LaViDa timing under single-block, EOS-suppressed decoding. N is the evaluation size; timing means are conditional on successful trace localization. U measures the rationale-side canvas, not a semantic rationale annotation. Coverage and accuracy use all evaluated outputs.
Figure 2: Recorded answer clocks and the block control. Markers are mean recorded times, and segments are differences of means. All rows are LaViDa with EOS suppressed; the short-block condition also changes answer-slot eligibility.
Data
Correct: mean U (%)
Incorrect: mean U (%)
p
M 3 CoT
98.5
97.1
0.68
ScienceQA
96.8
95.1
0.038
Table 2: Correctness-stratified LaViDa results. Mean U is the rationale-side fraction still unwritten at stabilization, expressed as a percentage. The last column gives the two-sided Mann–Whitney p -value for the correct/incorrect comparison.
Block
rpos
Tstab
Tcommit
U (%)
C (%)
A (%)
Q (%)
128
0.15
6.41
23.65
89.4
100.0
41.9
41.9
8
1.00
53.82
56.01
1.7
91.6
39.8
43.4
Table 3: Block-length control on LaViDa/V*Bench. Both rows evaluate 191 examples with the same configured canvas and step budget. Q is accuracy conditional on a valid extracted answer.
Figure 3: Visual-sensitivity timing. Markers are mean individual stabilization and PDM peak steps. Nrec denotes the recorded subset size. Each row gives the two means for one model–dataset configuration.
Model
Data
C
A
Q
Nemotron
V*Bench
90.6→99.5
47.6→52.4
52.6→52.6
Nemotron
M 3 CoT
73.5→88.8
54.5→69.5
74.1→78.3
Nemotron
ScienceQA
69.2→90.2
57.0→76.5
82.3→84.8
DiffusionGemma
V*Bench
84.3→91.1
53.9→58.6
64.0→64.4
LaViDa
V*Bench
84.8→66.5
39.8→28.8
46.9→43.3
Table 4: EOS-enabled prompting comparisons. Entries are reasoning-requested → direct (percent), with a shared scorer within each pair. C : valid-answer coverage; A : overall accuracy; Q : conditional accuracy. N=191 for V*Bench and 400 for the other rows. DiffusionGemma retains its own sampler.
Model
Data
ΔA
Coverage term
Conditional term
Nemotron
V*Bench
+4.8
+4.68
0.00
Nemotron
M 3 CoT
+15.0
+11.66
+3.41
Nemotron
ScienceQA
+19.5
+17.55
+1.99
DiffusionGemma
V*Bench
+4.7
+4.37
+0.35
LaViDa
V*Bench
−11.0
−8.25
−2.72
Table 5: Decomposition from the reported percentages. Values are percentage points. The two terms are computed from rounded C,Q and can sum slightly differently from the displayed ΔA ; the largest residual is 0.12 points.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Model
Data
Nrec
Mean Tstab
Mean TPDM
Final/peak (%)
LaViDa
V*Bench
191
6.3
3.6
25
LaViDa
M 3 CoT
146
0.7
0.8
13
LaViDa
ScienceQA
146
2.1
0.7
12
Nemotron
V*Bench
64
10.8
25.5
28
Nemotron
M 3 CoT
119
15.0
30.9
47
Nemotron
ScienceQA
118
14.0
30.5
45
Appendix
Table 6: PDM timing and relative final sensitivity. Nrec is the recorded subset size. The final/peak column is ρ from Equation 7 . Stabilization means belong to these PDM evaluations, not the primary timing conditions in Table 1 .
Diffusion vision-language models generate answers through iterative refinement, exposing intermediate answer trajectories that can be inspected and controlled at inference time. However, this controllability creates a reasoning-need mismatch, where a universal generation length is applied to questions with different reasoning demands. Visually closed questions may be harmed by continued refinement after a stable answer has formed, whereas reasoning-sensitive questions may be harmed by premature commitment. We formulate this problem as reasoning-budget mismatch and study it in LLaDA-V. Rather than choosing a universal generation length, our training-free controller routes each example to early commitment, baseline preservation, or reasoning-supportive decoding using trajectory signals from answer closure, commitment evidence, and representation revision pressure, without using ground-truth answers. Across answer-focused, mixed-reasoning, and CoT-sensitive benchmarks, routed control improves robustness over fixed long decoding, pure short decoding, and single-rule interventions. The gains are not explained by shorter outputs alone. Answer-closed examples often benefit from commitment, whereas CoT-sensitive examples require preserving or supporting intermediate reasoning. Taken together, these results suggest diffusion VLM decoding should route inference-time control by the state suggested by the observed trajectory instead of relying on a universal decoding length.
Yixiang Liu, Zhongxing Xu, Zhonghua Wang +1
Southern University of Science and Technology · Monash University
Masked diffusion language models (dLLMs) can commit tokens in any order -- a freedom marketed as their core advantage over autoregressive decoding. We show that on reasoning tasks this freedom is instead the axis of failure. Logging every commitment during decoding of LLaDA-8B on GSM8K, we find that unconstrained (pure) decoding commits the final answer at 15-24% of the trajectory while half the reasoning region is still masked, and collapses to answer-only outputs on up to 90% of problems as the canvas grows. The cause is not the model's termination beliefs -- EOS "pressure" is nearly identical across decoders -- but reachability: whether the sampler may act on those beliefs at distant positions. A 2x2 prompt-decoder design shows that chain-of-thought helps only under ordered commitment (interaction +34.8 percentage points, 95% CI [26.8, 42.8]; without reasoning text the decoders are indistinguishable), an interaction we decompose into a collapse channel and an order channel and replicate on Dream-7B and MATH-500. A single-knob intervention -- frontier-gated commitment -- causally recovers the full gap (0.528 to 0.852) while preserving up to 4x parallel decoding, along a measured frontier whose optimal window flips from w=1 at full refinement to unconstrained at 8 tokens/step. Our results reframe existing window-style samplers, previously motivated by efficiency, as the minimal fix for a reasoning pathology they were never designed to address.
Jewon Yeom, Jaewon Sok, Seonghyeon Park +3
Graduate School of Data Science, Seoul National University · Department of Rural Systems Engineering, Seoul National University · Department of Aerospace Engineering, Seoul National University
Vision Language Models (VLMs) achieve strong reasoning with Chain-of-Thought (CoT) prompting but incur high sequential-generation cost, error accumulation, and limited self-correction. Diffusion Multimodal Large Language Models (dMLLMs) unmask tokens in an order-agnostic process, improving efficiency and enabling iterative refinement, yet their reasoning and how to enhance it remain underexplored. We propose a training-free method, Spatio-Temporal Token Veto (ST-Veto), which leverages the ability to observe all token positions at each diffusion step. Rather than relying only on current-step confidence, ST-Veto vetoes temporally unstable tokens via second-order Taylor prediction of confidence dynamics and filters weakly grounded tokens using image-attention mass, swapping them with safer candidates. Across multiple dMLLMs and multimodal reasoning benchmarks, ST-Veto consistently outperforms standard decoding policies and prior VLM reasoning methods, improving accuracy by up to 9% with no additional training or generation cost. Analyses show that ST-Veto steers generation toward higher-confidence, better-grounded paths.
Keuntae Kim, Beomseok Lee, Hyunwoo Kim +1
Hanyang University · LG Electronics · KT Corporation