cs.CVOct 1, 2026

Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold

Authors: Keuntae Kim, Yong Suk Choi

Organizations: Department of Computer Science, Hanyang University, Seoul, Korea

Abstract

An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual question-answering benchmarks, we find that 89.4-98.1% of the rationale-side canvas remains unwritten at stabilization in single-block, EOS-suppressed LaViDa runs. On VBench, reducing block length from 128 to 8 changes this fraction from 89.4% to 1.7%, together with answer coverage and the eligible observation window. Under EOS-enabled prompting, direct instructions improve Nemotron's overall accuracy by 15.0 and 19.5 percentage points on M3CoT and ScienceQA, but reduce LaViDa/VBench accuracy by 11.0 points. A symmetric decomposition associates the larger absolute component of each change with coverage rather than conditional accuracy. Matched-canvas image ablations measure visual sensitivity alongside answer stabilization, separating the two temporal readouts. Together, these measurements distinguish answer stabilization, rationale unfolding, and visual sensitivity, and identify coverage as the larger component of the prompting differences.

Figures & tables

Appendix figures & tables1 asset

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models

    Sep 12, 2026Yixiang Liu, Zhongxing Xu, Zhonghua Wang +1Diffusion Language ModelsInference-Time Steering

  2. Answer First, Reason Later: Commitment Order in Diffusion LLMs

    Aug 6, 2026Jewon Yeom, Jaewon Sok, Seonghyeon Park +3Diffusion Language ModelsLLM Reasoning Strategies

  3. ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding

    Jul 20, 2026Keuntae Kim, Beomseok Lee, Hyunwoo Kim +1Diffusion Language ModelsMultimodal Reasoning Benchmarks