OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination
Organizations: Beijing University of Posts and Telecommunications · Nanyang Technological University, Singapore
Abstract
Omni-modal large language models (OmniLLMs) unify text, images, audio, and video, yet hallucinate when generation relies on the wrong evidence. Existing inference-time methods can reduce hallucinations, but rarely reveal which evidence sustains a generated commitment. We introduce OmniConfess, a training-free method for mitigating omni-modal hallucinations. It fixes a candidate response and re-scores it at token resolution under controlled channel-wise evidence interventions, producing a structured token-by-channel confession that reveals the response's evidential dependence. OmniConfess uses this confession to preserve grounded content and correct commitments driven by irrelevant or contradictory evidence. To evaluate OmniConfess, we construct OmniHalluBench, a 3,540-example benchmark built from six datasets spanning text, image, audio, and video settings and both judgment and free-form generation. Experiments show that OmniConfess mitigates hallucinations across heterogeneous modality and task settings. Our code and benchmark are publicly available at https://github.com/RongHuiQiang/OmniConfess.
Figures & tables
| Judgment tasks | Free-form tasks | Avg. | |||||||||||||
| PHD | CMM | PubMed | RAGTruth | HaloQ. | AVHBch | ||||||||||
| Method | F1 | Idx | F1 | Idx | F1 | Idx | F1 | GAV | F1 | GAV | F1 | GAV | F1 | Idx | GAV |
| Qwen2.5-Omni-7B | |||||||||||||||
| Base | 71.68 | 68.35 | 70.07 | 60.77 | 95.14 | 80.21 | 38.25 | 8.77 | 34.22 | 7.16 | 19.42 | 5.96 | 54.80 | 69.78 | 7.30 |
| SC | 76.43 | 73.18 | 71.96 | 71.31 | 91.25 | 68.29 | 35.38 | 8.89 | 34.48 | 7.14 | 18.29 | 6.04 | 54.63 | 70.93 | 7.36 |
| GD | 21.16 | 21.73 | 57.54 | 63.12 | 57.37 | 39.35 | 11.61 | 8.59 | 40.35 | 6.57 | 9.80 | 4.36 | 32.97 | 41.40 | 6.51 |
| Category | PHD | PubMedQA | CMM | RAGTruth | HaloQuest | AVHBench |
|---|---|---|---|---|---|---|
| Time · Tokens · Tok/s | Time · Tokens · Tok/s | Time · Tokens · Tok/s | Time · Tokens · Tok/s | Time · Tokens · Tok/s | Time · Tokens · Tok/s | |
| BaseModel | 2.23 · 23.4 · 10.5 | 0.81 · 25 · 30.9 | 20.64 · 22.7 · 1.1 | 2.69 · 98 · 36.4 | 0.93 · 10 · 10.8 | 2.65 · 25 · 9.4 |
| Search-based | 34.76 · 789.1 · 22.7 | 20.28 · 474.8 · 23.4 | 38.90 · 536.8 · 13.8 | 30.10 · 838.3 · 27.9 | 28.39 · 536.3 · 18.9 | 36.15 · 569.5 · 15.8 |
| Evidence-contrast | 7.17 · 21.5 · 3.0 | 1.19 · 26.3 · 22.1 | 25.44 · 22.9 · 0.9 | 2.49 · 69.4 · 27.9 | 1.32 · 10.5 · 8.0 | 4.09 · 22.1 · 5.4 |
| Fine-grained | 4.17 · 26.7 · 6.4 | 3.71 · 25.0 · 6.7 | 15.18 · 25.8 · 1.7 | 6.48 · 65.7 · 10.1 | 1.83 · 10.1 · 5.5 | 5.01 · 24.9 · 5.0 |
| OmniConfess | 5.43 · 29.3 · 5.4 | 1.23 · 25 · 20.3 | 35.14 · 24.6 · 0.7 | 4.82 · 89.2 · 18.5 | 0.82 · 12.5 · 15.2 | 7.93 · 164.1 · 20.7 |
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.