cs.AIApr 8, 2026

What Does a Sharing Question Add? Auditing LLM Survey Scores for Misinformation

Authors: Zonghuan Xu, Xiang Zheng, Yutao Wu, Xingjun Ma

Organizations: Shanghai Key Laboratory of Multimodal Embodied AI, Shanghai, China · City University of Hong Kong, Hong Kong SAR, China · Deakin University, Australia

Abstract

Evaluating misinformation requires distinguishing whether readers believe content from whether they would share it. Asking large language models (LLMs) both questions yields two scores, but does the sharing answer contribute information beyond the credibility answer? We audit eight model versions on 290 synthetic misinformation articles, using 1,256 paired survey responses with 317 participant identifiers as an external validity criterion. An initial reversal motivates the audit: every model's raw sharing score predicts mean human sharing less accurately than its credibility score. This ordering changes after offset correction, so it does not by itself diagnose missing information. We instead distinguish score reconstructability, persistence across elicitation formats, and incremental human validity. Credibility predicts 30.8-72.5% of model-sharing variation relative to a held-out constant baseline; remaining sharing differences correlate at 0.61-0.75 across question-order and separate-question conditions. Yet adding model sharing to human and model credibility yields only -0.25% to +0.69% error reduction with fixed regression, with all exploratory intervals crossing zero. Flexible prediction and format changes do not establish an improvement. Sharing answers therefore contain structured variation beyond the observed credibility score, without established incremental validity for human sharing in these data. The findings motivate validating the contribution of each elicited outcome, beyond inspecting score differences or agreement across prompts.

Figures & tables

Appendix figures & tables2 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility

    Aug 10, 2026Siyang Wu, Yibo Jiang, Bryon AragamLarge Language Model ReliabilityAnswer

  2. Large language models can effectively convince people to believe conspiracies

    Jan 8, 2026Thomas H. Costello, Kellin Pelrine, Matthew Kowal +6Large Language Model ReliabilityPersuasion

  3. VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation

    Sep 28, 2026Hanxun Huang, Yutao Wu, Qizhou Wang +9MisinformationFact-Checking