What Does a Sharing Question Add? Auditing LLM Survey Scores for Misinformation
Authors: Zonghuan Xu, Xiang Zheng, Yutao Wu, Xingjun Ma
Organizations: Shanghai Key Laboratory of Multimodal Embodied AI, Shanghai, China · City University of Hong Kong, Hong Kong SAR, China · Deakin University, Australia
Evaluating misinformation requires distinguishing whether readers believe content from whether they would share it. Asking large language models (LLMs) both questions yields two scores, but does the sharing answer contribute information beyond the credibility answer? We audit eight model versions on 290 synthetic misinformation articles, using 1,256 paired survey responses with 317 participant identifiers as an external validity criterion. An initial reversal motivates the audit: every model's raw sharing score predicts mean human sharing less accurately than its credibility score. This ordering changes after offset correction, so it does not by itself diagnose missing information. We instead distinguish score reconstructability, persistence across elicitation formats, and incremental human validity. Credibility predicts 30.8-72.5% of model-sharing variation relative to a held-out constant baseline; remaining sharing differences correlate at 0.61-0.75 across question-order and separate-question conditions. Yet adding model sharing to human and model credibility yields only -0.25% to +0.69% error reduction with fixed regression, with all exploratory intervals crossing zero. Flexible prediction and format changes do not establish an improvement. Sharing answers therefore contain structured variation beyond the observed credibility score, without established incremental validity for human sharing in these data. The findings motivate validating the contribution of each elicited outcome, beyond inspecting score differences or agreement across prompts.
Figures & tables
Component
Retained observations
Content
290 articles, 99 claim–goal scenarios
Human criterion
1,256 paired records, 317 study IDs
Main model scores
8 versions × 290 score pairs
Format comparisons
Claude: 290 articles; GPT-5.4 and Gemini 3.1 Pro: 99 each
Formats
C then S ; S then C ; separate calls
Response scale
Credibility C and sharing S : 1–7
Table 1. Audit scope. Model scores vary by article, not by human participant. Each listed model condition retains one successful response per article and item.
Figure 1. Motivating article-mean diagnostic. Lower is better; the dashed line is the training-mean baseline. Raw S is worse than C for all eight evaluators, whereas correcting the offset reverses that ordering. Affine calibration brings both near the constant. Each mapping is fitted with scenario-held-out validation. Axes have different ranges to show the successive changes. This diagnostic evaluates numerical use of scores, not the conditional human-validity estimand in Equation 4 . Three panels compare credibility and sharing error before calibration, after offset correction, and after affine calibration across eight model versions.
Figure 2. Reconstructability and external validity answer different questions. (a) Held-out reconstruction of a model’s sharing score from its own credibility score, relative to a training constant. (b) Increment from adding model sharing when predicting individual human sharing, conditional on human credibility, model credibility, and wave. Fixed and training-selected fits use scenario holdout with test-participant exclusion. Bars are exploratory 95% intervals: scenario resampling in (a), crossed participant–scenario resampling in (b). The two panels have different targets and denominators. Credibility reconstructs some but not all model sharing variation. The human-validity gains cluster near zero and all intervals cross zero.
Figure 3. Cross-format correspondence of sharing residuals after each condition’s observed credibility categories are removed. Each comparison uses the original C -then- S condition as the reference. Claude has 290 articles; the other models have 99. Intervals resample scenarios while holding residuals fixed. This is not same-condition test–retest reliability. Sharing residuals correlate positively across reversed order and separate-question conditions in three model versions.
Evaluator
Elicitation format
Articles
Reconstruction (%)
Human increment (%)
Claude 4.5
C then S
290
55.8
-0.02 [-0.70, +0.69]
Claude 4.5
S then C
290
53.0
+0.18 [-0.65, +1.10]
Claude 4.5
Separate
290
24.0
+0.12 [-0.71, +0.99]
GPT-5.4
C then S
99
28.0
+0.17 [-2.12, +2.05]
GPT-5.4
S then C
99
42.0
+0.36 [-2.44, +3.05]
GPT-5.4
Separate
99
27.2
-1.22 [-4.64, +0.83]
Table 2. Format-specific results on matched article sets. Reconstruction is the percentage reduction in model-sharing MSE relative to a constant, using nested selection. Human increment is the fixed-fit relative reduction from adding model sharing to human and model credibility and wave. Brackets give exploratory 95% crossed-resampling intervals. Smaller reconstruction is not itself better: it must be interpreted alongside external validity.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
Evaluator
Reconstruction (%)
Residual RMSE
Fixed increment (%)
Selected increment (%)
Claude 4.5
63.4
0.439
-0.06 [-0.94, +0.86]
-0.05 [-0.15, +0.02]
GPT-4.1
54.3
0.819
-0.18 [-1.48, +1.06]
-0.58 [-3.01, +1.50]
GPT-4o
72.5
0.667
+0.33 [-1.40, +2.01]
-0.60 [-2.49, +0.95]
GPT-5
48.5
0.727
-0.25 [-1.07, +0.46]
+0.09 [-0.25, +0.61]
GPT-5.1
37.3
0.898
+0.69 [-1.67, +2.71]
+0.44 [-1.76, +2.43]
GPT-5.2
32.7
0.640
+0.30 [-1.14, +1.71]
+0.33 [-0.85, +1.45]
Appendix
Table 3. Main per-model results. Reconstruction uses nested selection; its RMSE is in score points. The last two columns report human-sharing MSE reductions from adding model sharing to human and model credibility and wave. Brackets are exploratory 95% crossed-resampling intervals.
Evaluator
Model-only sharing increment (%)
Reverse credibility increment (%)
Claude 4.5
+0.03 [-0.49, +0.56]
+2.03 [-1.90, +6.70]
GPT-4.1
-0.04 [-1.91, +1.66]
+2.82 [-2.44, +8.05]
GPT-4o
+0.23 [-0.26, +0.80]
+3.12 [-2.33, +8.71]
GPT-5
+0.01 [-0.14, +0.17]
+2.28 [-2.54, +6.84]
GPT-5.1
+0.20 [-0.57, +0.94]
+4.64 [-1.42, +10.35]
GPT-5.2
+0.08 [-0.64, +0.77]
+4.64 [-1.03, +10.24]
Appendix
Table 4. Additional human-validity comparisons. Left: adding model sharing to model credibility and wave, without human credibility (selected fit). Right: adding model credibility to human sharing, model sharing, and wave to predict human credibility (fixed fit). Each column has its own baseline and target; differences between columns are not tests of construct superiority.
Data Science Institute University of Chicago Chicago, USA · Department of Computer Science University of Chicago Chicago, USA · Booth School of Business University of Chicago Chicago, USA