Organizations: Nanjing University of Aeronautics and Astronautics · State Key Laboratory of Ocean Sensing, ZJU-Hangzhou Global Scientific and Technological Innovation Center, Zhejiang University, Hangzhou, 311215, China · Security Capability Center
Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsupported claim in diffusion VLMs appears late or has already stabilized before any answer token is revealed. We introduce DynaHall, a trajectory-level benchmark of annotation-backed binary visual propositions covering object existence, counting, attributes, and relations, with controlled hard negatives graded by visual prior. DynaHall is paired with a commitment-aware protocol that records the intermediate answer tendency at every unmasking step alongside the committed output. Across five diffusion VLMs from three architecture families, visual hallucination is settled before commitment: an unsupported answer is already the preferred state while the answer position is still masked, and later unmasking steps rarely reverse it, so the failure is not introduced at the write step. This holds across decoding schedules, answer formats, and open-ended generation. DynaHall also exposes failures hidden by final-output metrics, including counting and relation collapse, prior-driven false positives, and attribute errors whose direction changes by type. Guided by this diagnosis, PGS (Pre-commitment Gradient Steering) edits still-masked answer states to reduce false positives, bringing the affirmation rate close to balance, and transfers to another architecture without degrading general ability. DynaHall and PGS suggest that hallucination should be measured and mitigated along the generation trajectory of diffusion VLMs, not only at the final answer.
Figures & tables
Figure 1: Final-output benchmarks observe only the committed endpoint, whereas DynaHall tracks the full unmasking trajectory. The example shows a hallucinated answer tendency that is already stable before the answer token is committed.
Target
Evaluation design
Method
Benchmark / study
Visual
Diffusion
Trajectory
Graded neg.
Bias-ctrl.
Annot.
Mitigation
LVLM hallucination benchmarks (autoregressive)
POPE ( Li et al. 2023 )
✓
×
×
✓
✓
✓
×
AMBER ( Wang et al. 2023 )
✓
×
×
×
∘
✓
×
HallusionBench ( Guan et al. 2024 )
✓
×
×
∘
✓
×
×
Reefknot ( Zheng et al. 2025 )
✓
×
×
∘
∘
✓
✓
Table 1: Positioning of DynaHall against representative hallucination benchmarks and diffusion-LM trajectory studies. ✓/ ∘ / × denote yes/partial/no; † ( Hemmat et al. 2026 ; Qian et al. 2026 ; Zhao et al. 2026 ) .
Figure 2: Construction of DynaHall . From COCO instances and GQA scene graphs (a), we build binary visual propositions for four tasks (b), each paired with a controlled, prior-graded hard negative (c); both polarities form a balanced set of 8,000 propositions (d) validated by expert audit (e). A single tennis scene runs through every stage.
Task
#Pos
#Neg
Total
Object existence
1,000
1,000
2,000
Counting
1,000
1,000
2,000
Attribute
1,000
1,000
2,000
Relation
1,000
1,000
2,000
Total
4,000
4,000
8,000
Construction summary
Table 2: Dataset statistics for DynaHall . The benchmark is balanced by task and answer label.
Figure 3: Trajectory readout for one running example (“Is the man holding the cap?”, gold no): the answer tendency P(yes)/P(no) at the locus p across unmasking steps, with the answer token committed only at the final step.
Final output
Trajectory dynamics ( DynaHall )
Model
Family
Acc ↑
FP ↓
FN ↓
Aff.
Early-FP
Emerg.
Persist.
Correct.
Drift
Flips
MMaDA-Base
MMaDA/VQ
66.2
42.5
25.0
58.7
96.5
0.14
28.0
21.8
9.0
1.08
MMaDA-MixCOT
MMaDA/VQ
74.1
14.5
37.5
38.4
78.8
0.15
20.9
19.9
6.7
0.99
Lumina-DiMOO
VQ-diff.
73.6
6.4
46.6
29.8
92.0
0.06
24.7
8.6
3.5
0.23
LLaDA-V ∗
cont.-img
75.9
19.7
13.8
46.8
76.9
0.08
15.9
27.7
12.0
0.50
Dimple-7B
Qwen-diff.
81.4
19.8
17.5
51.0
91.8
0.14
16.3
22.7
4.7
0.19
Table 3: Model-level hallucination diagnosis on the DynaHall test set (percentages). Final-output columns are standard metrics; the shaded Trajectory-dynamics columns are unique to DynaHall and describe how errors evolve, with no preferred direction (Emergence is a normalized step, Flips a count). Best per oriented column ( ↑ / ↓ ) in bold . Correction/Drift are the shares of ever-wrong/ever-correct trajectories that recover or later flip; Persistence is the share of all trajectories that remain wrong continuously from their first erroneous tendency through the final output. ∗ LLaDA-V’s 7.3% invalid outputs count as accuracy errors and as non-affirmative for Aff.; they enter neither the FP nor FN numerator.
Figure 4: Stage-wise false-positive rate on negative propositions across unmasking for five diffusion VLMs ( DynaHall test). Rates are read from the tendency state (logit preference at the masked locus) and can therefore differ from the committed FP in Table 3 . Shaded bands are bootstrap 95% confidence intervals; the inset reports each model’s FP slope across the unmasking process.
Model
Attr.
Acc ↑
FP ↓
FN ↓
MMaDA-Base
color
72.2
20.1
35.2
material
65.0
9.0
61.0
shape
60.3
30.5
49.0
Dimple-7B
color
85.8
7.8
20.2
material
86.5
2.0
25.0
shape
77.5
14.5
30.5
Table 4: Attribute-subtype diagnosis for two representative models, on DynaHall test. Columns are accuracy, FP on negatives, and FN on positives (percentages).
Diffusion-based Large Vision-Language Models (dLVLMs) have recently emerged as a compelling alternative to autoregressive (AR) LVLMs, offering advantages in parallel decoding, bidirectional context, and controllable generation. Despite rapid progress, their reliability properties remain largely uncharacterized. We present the first systematic reliability evaluation of hallucination and bias in dLVLMs, benchmarking six diffusion models against competitive AR baselines across four dimensions. Our key findings are: (1) dLVLMs reverse the yes-bias of AR models in binary visual queries; (2) they achieve competitive hallucination rates yet exhibit degraded linguistic quality; (3) they collapse to near-zero accuracy on underrepresented racial groups with opposite-polarity gender bias; and (4) they exhibit accuracy collapse in multiple-choice settings when the correct option is shorter than its distractors, associated with a length prior that emerges at the first denoising step. Tokens committed at late denoising steps with low confidence further correlate with hallucinated content, pointing to a mechanistic signal unique to diffusion generation. These patterns vary across model families, suggesting reliability is shaped by the generative paradigm together with training data.
Diffusion large language models generate text through iterative denoising, exposing hidden trajectories that may contain reliability signals beyond the final output. We propose HIVE, which compresses trajectory hidden states, selects informative step-layer evidence, and conditions a verifier through continuous prefix embeddings to produce a hallucination score and structured diagnostics. Across two D-LLMs and three QA benchmarks, HIVE outperforms eight established baselines and a verifier-backbone-matched text-only control in all six settings. Relative to text-only verification, hidden-evidence conditioning improves AUROC by 1.73--4.60 points and AUPRC by 1.10--3.62 points, with average gains of 3.15 and 2.28 points, respectively. Ablations, evidence interventions, and cross-dataset transfer further support the complementary value of fine-grained hidden trajectory evidence.
Guoshenghui Zhao, Tan Yu, Weijie Zhao
Rochester Institute of Technology · NVIDIA Corporation
Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the visual input. Prior work has attributed hallucinations in LVLMs to factors such as limitations of the vision backbone or the dominance of the language component, yet the relative importance of these factors remains unclear. To resolve this ambiguity, We propose HalluScope, a benchmark to better understand the extent to which different factors induce hallucinations. Our analysis indicates that hallucinations largely stem from excessive reliance on textual priors and background knowledge, especially information introduced through textual instructions. To mitigate hallucinations induced by textual instruction priors, we propose HalluVL-DPO, a framework for fine-tuning off-the-shelf LVLMs towards more visually grounded responses. HalluVL-DPO leverages preference optimization using a curated training dataset that we construct, guiding the model to prefer grounded responses over hallucinated ones. We demonstrate that our optimized model effectively mitigates the targeted hallucination failure mode, while preserving or improving performance on other hallucination benchmarks and visual capability evaluations. To support reproducibility and further research, we will publicly release our evaluation benchmark, preference training dataset, and code at https://pegah-kh.github.io/projects/prompts-override-vision/ .
Pegah Khayatan, Jayneel Parekh, Arnaud Dapogny +3
1ISIR, Sorbonne Universit´e, Paris, France · 2Valeo.ai, Paris, France