Organizations: Nanjing University of Aeronautics and Astronautics · State Key Laboratory of Ocean Sensing, ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University, Hangzhou, 311215, China · Security Capability Center
Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsupported claim in diffusion VLMs appears late or has already stabilized before any answer token is revealed. We introduce DynaHall, a trajectory-level benchmark of annotation-backed binary visual propositions covering object existence, counting, attributes, and relations, with controlled hard negatives graded by visual prior. DynaHall is paired with a commitment-aware protocol that records the intermediate answer tendency at every unmasking step alongside the committed output. Across five diffusion VLMs from three architecture families, visual hallucination is settled before commitment: an unsupported answer is already the preferred state while the answer position is still masked, and later unmasking steps rarely reverse it, so the failure is not introduced at the write step. This holds across decoding schedules, answer formats, and open-ended generation. DynaHall also exposes failures hidden by final-output metrics, including counting and relation collapse, prior-driven false positives, and attribute errors whose direction changes by type. Guided by this diagnosis, PGS (Pre-commitment Gradient Steering) edits still-masked answer states to reduce false positives, bringing the affirmation rate close to balance, and transfers to another architecture without degrading general ability. DynaHall and PGS suggest that hallucination should be measured and mitigated along the generation trajectory of diffusion VLMs, not only at the final answer.
Figures & tables
Figure 1: Final-output benchmarks observe only the committed endpoint, whereas DynaHall tracks the full unmasking trajectory. The example shows a hallucinated answer tendency that is already stable before the answer token is committed.
Target
Evaluation design
Method
Benchmark / study
Visual
Diffusion
Trajectory
Graded neg.
Bias-ctrl.
Annot.
Mitigation
LVLM hallucination benchmarks (autoregressive)
POPE ( Li et al. 2023 )
✓
×
×
✓
✓
✓
×
AMBER ( Wang et al. 2023 )
✓
×
×
×
∘
✓
×
HallusionBench ( Guan et al. 2024 )
✓
×
×
∘
✓
×
×
Reefknot ( Zheng et al. 2025 )
✓
×
×
∘
∘
✓
✓
Table 1: Positioning of DynaHall against representative hallucination benchmarks and diffusion-LM trajectory studies. ✓/ ∘ / × denote yes/partial/no; † ( Hemmat et al. 2026 ; Qian et al. 2026 ; Zhao et al. 2026 ) .
Figure 2: Construction of DynaHall . From COCO instances and GQA scene graphs (a), we build binary visual propositions for four tasks (b), each paired with a controlled, prior-graded hard negative (c); both polarities form a balanced set of 8,000 propositions (d) validated by expert audit (e). A single tennis scene runs through every stage.
Task
#Pos
#Neg
Total
Object existence
1,000
1,000
2,000
Counting
1,000
1,000
2,000
Attribute
1,000
1,000
2,000
Relation
1,000
1,000
2,000
Total
4,000
4,000
8,000
Construction summary
Table 2: Dataset statistics for DynaHall . The benchmark is balanced by task and answer label.
Figure 3: Trajectory readout for one running example (“Is the man holding the cap?”, gold no): the answer tendency P(yes)/P(no) at the locus p across unmasking steps, with the answer token committed only at the final step.
Final output
Trajectory dynamics ( DynaHall )
Model
Family
Acc ↑
FP ↓
FN ↓
Aff.
Early-FP
Emerg.
Persist.
Correct.
Drift
Flips
MMaDA-Base
MMaDA/VQ
66.2
42.5
25.0
58.7
96.5
0.14
28.0
21.8
9.0
1.08
MMaDA-MixCOT
MMaDA/VQ
74.1
14.5
37.5
38.4
78.8
0.15
20.9
19.9
6.7
0.99
Lumina-DiMOO
VQ-diff.
73.6
6.4
46.6
29.8
92.0
0.06
24.7
8.6
3.5
0.23
LLaDA-V ∗
cont.-img
75.9
19.7
13.8
46.8
76.9
0.08
15.9
27.7
12.0
0.50
Dimple-7B
Qwen-diff.
81.4
19.8
17.5
51.0
91.8
0.14
16.3
22.7
4.7
0.19
Table 3: Model-level hallucination diagnosis on the DynaHall test set (percentages). Final-output columns are standard metrics; the shaded Trajectory-dynamics columns are unique to DynaHall and describe how errors evolve, with no preferred direction (Emergence is a normalized step, Flips a count). Best per oriented column ( ↑ / ↓ ) in bold . Correction/Drift are the shares of ever-wrong/ever-correct trajectories that recover or later flip; Persistence is the share of all trajectories that remain wrong continuously from their first erroneous tendency through the final output. ∗ LLaDA-V’s 7.3% invalid outputs count as accuracy errors and as non-affirmative for Aff.; they enter neither the FP nor FN numerator.
Figure 4: Stage-wise false-positive rate on negative propositions across unmasking for five diffusion VLMs ( DynaHall test). Rates are read from the tendency state (logit preference at the masked locus) and can therefore differ from the committed FP in Table 3 . Shaded bands are bootstrap 95% confidence intervals; the inset reports each model’s FP slope across the unmasking process.
Model
Attr.
Acc ↑
FP ↓
FN ↓
MMaDA-Base
color
72.2
20.1
35.2
material
65.0
9.0
61.0
shape
60.3
30.5
49.0
Dimple-7B
color
85.8
7.8
20.2
material
86.5
2.0
25.0
shape
77.5
14.5
30.5
Table 4: Attribute-subtype diagnosis for two representative models, on DynaHall test. Columns are accuracy, FP on negatives, and FN on positives (percentages).