Examining Social Attribution in LLM Reasoning: A Theory-Guided Probing Methodology
Authors: Zhaoxin Yu, Qingchao Kong, Dajun Zeng, Wenji Mao
Organizations: State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences
Large language models (LLMs) are increasingly deployed in sociotechnical systems where social attribution, the reasoning process attributing external events to the causes and reasons of agents' social behaviors, plays a critical role. These processes involve judgments of social cause, responsibility, and blame/credit to agents. Although attributional models are well-studied in social psychology and cognition through Attribution Theory, social attribution remains underexplored in AI, particularly LLM social reasoning. This paper provides the first systematic exploration of LLM social attribution. Our work focuses on responsibility and blame attributions, examining current LLMs' judgments and their underlying internal mechanisms. Guided by attribution theory, we construct a social attribution benchmark consisting of a Vignette subset based on classic scenarios from attribution theory research and a Reality subset based on real-world social narratives, yielding 7,639 responsibility/blame judgment questions. On this basis, we evaluate 32 representative LLMs and 5 basic non-LLM baselines. To further explore the internal mechanisms underlying the LLM judgment process, we develop a probing-based methodology to investigate the latent-space representations of 5 key attribution dimensions and the consistency of their influences on LLM judgments compared to those in human social attribution. Our research findings reveal that current LLMs exhibit measurable but incomplete agreement with human responsibility and blame judgments, and meanwhile, this agreement is positively correlated with model size. Some attribution dimensions are systematically decodable from specific positions in LLM hidden states, and their influences on the final judgment are consistent with those indicated by human Attribution Theory. The dataset and associated code are available at https://github.com/Yuzhaoxin946/SAB-Bench.
Figures & tables
Figure 1: Benchmark construction and evaluation of LLM social attribution.
Figure 2: Attribution-dimension probing through layerwise decodability analysis and directional interventions on judgments.
Figure 3: Accuracy and Spearman’s ρ across models, with 95% confidence intervals.
Figure 4: Model-size trends averaged over both tasks, with OLS fits and 95% confidence bands.
Figure 5: Layerwise dimension decodability. Panels separate models and tasks; rows denote dimensions, and darker colors indicate higher decodability (chance: 50%).
Figure 6: Layerwise influence of attribution dimensions. Blue cells pass all three question-level directional-success tests ( p<0.01 ); gray cells do not. Rows denote dimensions; 100% denotes the final layer.
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Classic attribution theories organize responsibility and blame judgments around assessments of causality, mental states, control, and mitigating circumstances ( Weiner, 1995 ; Shaver, 1985 ; Malle et al., 2014 ) . The models provide complementary theoretical foundations for the dimensions investigated here; they are not assumed to describe the computational architecture of an LLM.
Table 1: The 11 Vignette templates span eight scenario types and three inter-agent relation categories. Construction settings precede annotation filtering. “Agents Number” gives the number of agents in each scenario, and “#Instances” gives the number of scenario instances per language. Fixed and variable dimensions are subject to agent-specific role constraints.
Type
Task
κ
κlin
κquad
EN
ZH
EN
ZH
EN
ZH
General
Resp.
90.95
91.53
93.89
94.38
96.50
96.88
Blame
84.72
85.42
90.52
91.03
95.01
95.35
Vicarious blame
Resp.
77.18
81.69
79.85
83.63
83.67
86.50
Blame
74.49
78.88
81.93
85.51
88.84
91.57
Commanding chain
Resp.
80.86
84.86
85.95
88.85
90.97
92.76
Appendix
Table 2: Inter-annotator agreement ( ×100 ). Resp. denotes responsibility. All coefficients are significantly different from zero ( p<0.001 ).
Subset
Before
After
Removed
Vignette
5,344 (60.7%)
4,776 (62.5%)
568
Reality
3,456 (39.3%)
2,863 (37.5%)
593
Total
8,800
7,639
1,161
Appendix
Table 3: Question counts before and after agreement-based filtering. Percentages give each subset’s share within the Before or After column; both tasks and languages are pooled.
Agent relation
Before
After
General Scenarios
3,840 (71.9%)
3,488 (73.0%)
Vicarious Blame Scenarios
640 (12.0%)
541 (11.3%)
Commanding Chain Scenarios
864 (16.2%)
747 (15.7%)
Appendix
Table 4: Distribution of the three Vignette relation categories before and after filtering. Percentages use the Vignette total at the corresponding stage.
Dimension
Before
After
Value = 1
Value = 0
Value = 1
Value = 0
Intention
6,097 (69.3%)
2,703 (30.7%)
5,405 (70.8%)
2,234 (29.2%)
Voluntariness
3,164 (67.7%)
1,509 (32.3%)
2,845 (69.2%)
1,267 (30.8%)
Foreknowledge
4,572 (52.0%)
4,228 (48.0%)
3,935 (51.5%)
3,704 (48.5%)
Controllability
5,462 (62.1%)
3,338 (37.9%)
4,709 (61.6%)
2,930 (38.4%)
Obligation
3,087 (50.8%)
2,993 (49.2%)
2,740 (53.2%)
2,413 (46.8%)
Appendix
Table 5: Attribution-dimension distributions before and after filtering. Percentages use questions with non-missing labels for each dimension at the corresponding stage.
Subset
Task
Stage
High
Medium
Low
No
Total
Vignette
Responsibility
Before
961 (36.0%)
919 (34.4%)
598 (22.4%)
194 (7.3%)
2,672
After
921 (37.8%)
791 (32.4%)
539 (22.1%)
188 (7.7%)
2,439
Blame
Before
583 (21.8%)
733 (27.4%)
744 (27.8%)
612 (22.9%)
2,672
After
556 (23.8%)
550 (23.5%)
639 (27.3%)
592 (25.3%)
2,337
Reality
Responsibility
Before
815 (47.2%)
528 (30.6%)
189 (10.9%)
196 (11.3%)
1,728
After
730 (51.2%)
369 (25.9%)
146 (10.2%)
181 (12.7%)
1,426
Appendix
Table 6: Judgment-label distributions before and after agreement-based filtering, pooled across English and Chinese. Percentages use the row total.
Subset
Task
Language
Judgment
EN
ZH
High
Medium
Low
No
Vignette
Resp.
1,226 (50.3%)
1,213 (49.7%)
921 (37.8%)
791 (32.4%)
539 (22.1%)
188 (7.7%)
Vignette
Blame
1,171 (50.1%)
1,166 (49.9%)
556 (23.8%)
550 (23.5%)
639 (27.3%)
592 (25.3%)
Reality
Resp.
720 (50.5%)
706 (49.5%)
730 (51.2%)
369 (25.9%)
146 (10.2%)
181 (12.7%)
Reality
Blame
720 (50.1%)
717 (49.9%)
634 (44.1%)
321 (22.3%)
183 (12.7%)
299 (20.8%)
Appendix
Table 7: Language and judgment distributions. Each cell reports a count and its percentage within the row. Resp. denotes responsibility.
Dimension
Value 1
Value 0
Intention
5,405 (70.8%)
2,234 (29.2%)
Voluntariness
2,845 (69.2%)
1,267 (30.8%)
Foreknowledge
3,935 (51.5%)
3,704 (48.5%)
Controllability
4,709 (61.6%)
2,930 (38.4%)
Obligation
2,740 (53.2%)
2,413 (46.8%)
Appendix
Table 8: Attribution-dimension distributions. Percentages are calculated within each dimension using only questions with non-missing labels.
Agent Relation
Scenario Type
Count
Percentage
General
Traffic Accident
755
15.8%
General
Food Allergy
744
15.6%
General
Company Scenario
759
15.9%
General
Drowning Incident
764
16.0%
General
Laboratory Scenario
466
9.8%
Vicarious Blame
Company Scenario
248
5.2%
Appendix
Table 9: Retained Vignette questions by agent relation and scenario type. The 11 rows cover eight distinct scenario types, with shared types repeated under their respective relations. Percentages use all 4,776 Vignette questions as the denominator.
Figure 8: High-frequency content words in the English Reality narratives.
Figure 9: High-frequency content words in the Chinese Reality narratives.
LLM Family
Developer
Open/Closed-source
Version
Size
Claude
Anthropic
Closed
Claude-Haiku-4.5
-
Doubao Seed
ByteDance
Closed
Doubao-Seed-1.8
-
Gemini
Google
Closed
Gemini-3.5-flash
-
ChatGPT
OpenAI
Closed
GPT-5-mini / GPT-5-nano
-
Grok
xAI
Closed
Grok-4.3
-
Kimi
Moonshot AI
Closed
Kimi-K2.5
-
Appendix
Table 10: List of LLMs under test and their detailed information.
Backbone
L
20%
40%
60%
80%
Final
DeepSeek-R1-1.5B
28
5
11
16
22
28
Llama-3.2-1B
16
3
6
9
12
16
Llama-3.1-8B
32
6
12
19
25
32
Qwen3-4B
36
7
14
21
28
36
Qwen3-8B
36
7
14
21
28
36
Appendix
Table 11: Backbones and intervention layers. Layer numbers use the one-based convention l∈{1,…,L} . The four intermediate positions are ⌊0.2L⌋ , ⌊0.4L⌋ , ⌊0.6L⌋ , and ⌊0.8L⌋ ; percentages are nominal depth selections.
Setting
Value or procedure
Prompt strategy
CoT
Probe-training runs
Separate for each language and subset
Vignette cross-validation
8 folds; one scenario type held out, seven used for training
Reality cross-validation
10 random folds; one held out, nine used for training
Probe-training seeds
0, 1, 2
Decodability score
Balanced accuracy; chance 50%
Appendix
Table 12: Settings shared across the probing and intervention analyses. Backbone-specific layer selections appear in Table 11 .
Prompt Strategy
Model
N
CoT
SR
TG
Avg.
95% CI
Closed-source Models
Claude-Haiku-4.5
55.71
52.01
54.75
56.09
54.64
[53.16, 56.16]
Doubao-Seed-1.8
52.70
52.91
52.39
54.51
53.13
[51.34, 54.81]
Gemini-3.5-flash
50.09
52.19
52.82
52.73
51.96
[50.35, 53.77]
GPT-5-mini
57.18
57.67
55.96
56.66
56.87
[55.36, 58.38]
Appendix
Table 13: Accuracy of 32 representative LLMs on our constructed social attribution benchmark.
Prompt Strategy
Model
N
CoT
SR
TG
Avg.
95% CI
Closed-source Models
Claude-Haiku-4.5
0.554
0.534
0.538
0.544
0.542
[0.509, 0.575]
Doubao-Seed-1.8
0.535
0.540
0.530
0.532
0.534
[0.495, 0.571]
Gemini-3.5-flash
0.519
0.554
0.521
0.541
0.534
[0.497, 0.570]
GPT-5-mini
0.556
0.567
0.538
0.534
0.549
[0.515, 0.584]
Appendix
Table 14: Spearman ρ of 32 representative LLMs on our constructed social attribution benchmark.
Baseline
Acc.
95% CI
ρ
95% CI
Random Stratified
30.16
[28.62, 31.68]
0.028
[-0.013, 0.068]
Majority Class
40.93
[38.87, 42.99]
0.000
[0.000, 0.000]
Heuristic Rule
21.51
[19.82, 23.16]
0.115
[0.078, 0.149]
Attribution Rule
33.52
[31.52, 35.43]
0.303
[0.246, 0.353]
mBERT Classifier
43.47
[41.43, 45.32]
0.141
[0.112, 0.170]
Appendix
Table 15: Social attribution performance of basic baselines.
Prompt Strategy
Model
N
CoT
SR
TG
Avg.
Closed-source Models
Claude-Haiku-4.5
100.00
100.00
100.00
100.00
100.00
Doubao-Seed-1.8
100.00
100.00
100.00
100.00
100.00
Gemini-3.5-flash
99.90
100.00
100.00
100.00
99.97
GPT-5-mini
100.00
100.00
100.00
100.00
100.00
Appendix
Table 16: Valid output rates of evaluated LLMs by task and prompt strategy.
Vignette agent relation
Model
General
Vicarious
Commanding
Reality
Closed-source Models
Claude-Haiku-4.5
55.74
42.92
41.44
58.98
Doubao-Seed-1.8
48.83
44.62
39.15
63.83
Gemini-3.5-flash
46.60
48.84
43.60
61.46
GPT-5-mini
55.87
51.88
46.16
61.89
Appendix
Table 17: Accuracy by inter-agent relation within Vignette , with Reality shown separately.
Vignette agent relation
Model
General
Vicarious
Commanding
Closed-source Models
Claude-Haiku-4.5
0.661
0.256
0.258
Doubao-Seed-1.8
0.646
0.252
0.275
Gemini-3.5-flash
0.627
0.324
0.299
GPT-5-mini
0.645
0.334
0.303
Appendix
Table 18: Spearman ρ by inter-agent relation within Vignette .
(a) Responsibility judgment
(b) Blame judgment
Intent.
Volunt.
Foreknow.
Control.
Obligation
Intent.
Volunt.
Foreknow.
Control.
Obligation
Model
0
1
0
1
0
1
0
1
0
1
0
1
0
1
0
1
0
1
0
1
Closed-source Models
Claude-Haiku-4.5
44.55
58.68
53.42
66.64
46.43
62.27
48.44
58.47
56.89
55.00
45.58
53.46
49.73
54.69
43.45
58.55
46.02
54.62
56.23
51.69
Doubao-Seed-1.8
41.92
57.62
54.33
67.43
43.53
62.06
40.53
60.93
56.84
52.10
49.60
54.40
57.48
57.03
45.75
59.83
47.69
56.26
55.51
51.88
Gemini-3.5-flash
41.10
56.30
52.86
63.98
40.59
62.52
41.40
58.49
54.46
53.86
49.98
55.09
54.00
58.12
50.96
56.04
50.15
55.70
55.45
54.17
Appendix
Table 19: Accuracy of tested models by attribution dimension values.
Figure 10: Sanity-check results at the label-token prediction step for DeepSeek-R1-1.5B, including shuffled-label, 64-dimensional random-projection, scenario-text lexical, and chance controls.
Figure 11: Sanity-check results at the label-token prediction step for Llama-3.2-1B, including shuffled-label, 64-dimensional random-projection, scenario-text lexical, and chance controls.
Figure 12: Sanity-check results at the label-token prediction step for Llama-3.1-8B, including shuffled-label, 64-dimensional random-projection, scenario-text lexical, and chance controls.
Figure 13: Sanity-check results at the label-token prediction step for Qwen3-4B, including shuffled-label, 64-dimensional random-projection, scenario-text lexical, and chance controls.
Figure 14: Sanity-check results at the label-token prediction step for Qwen3-8B, including shuffled-label, 64-dimensional random-projection, scenario-text lexical, and chance controls.
Figure 15: Layerwise decodability at the last input token and the label-token prediction step. Layer depth is normalized within each model; dots indicate peak decodability.
Figure 16: Layerwise influence of attribution dimensions when interventions are applied at the last input token, with α=8 . Blue cells pass all three question-level directional-success tests ( p<0.01 ); gray cells do not. Rows denote dimensions; 100% denotes the final layer. Int., Vol., For., Con., and Obl. denote intention, voluntariness, foreknowledge, controllability, and obligation.
Figure 17: Average signed judgment change D(α) at different intervention strengths, shown separately for responsibility and blame. Each model’s curve averages Dl,k(α) equally across evaluated layers and dimensions. Positive values indicate a net theory-aligned change, not a directional-success rate or a significance result.
We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social reasoning versus STEM reasoning in OLMo3-7B. Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities. We compute gradient-based attribution (TrackStar via Bergson) over a working set drawn from the de-duplicated Dolma3 mix, aggregate influence across WebOrganizer's 24-format x 24-topic taxonomy (576 bins), and contrast benchmark pairs in a 2x2 design that varies domain (social vs. STEM) and capability type (reasoning vs. knowledge): SocialIQA and MMLU Social Sciences against ARC-Challenge and MMLU STEM. Social and STEM reasoning draw on qualitatively distinct corpus regions, and the contrast is sharper at the reasoning level than at the knowledge level. Targeted machine unlearning provides partial causal validation: forgetting high-attribution topics (e.g., Literature for SocialIQA) degrades the aligned benchmark more than within-topic random baselines. We release the code and aggregate artifacts at https://github.com/HCAI-Lab-GT/capabilibara and https://huggingface.co/HCAI-Lab-GT.
The generative nature of Large Language Models (LLMs) is reflected in the conditional probabilities they compute to sample each response token given the previous tokens. These probabilities encode the distributional structure that the model learns in training and exploits in inference. In this work, we use these probabilities to situate LLMs within the mathematical theory of stochastic processes. We use this framework to design a model-agnostic probabilistic token attribution measure, using Bayes rule to invert the next-token log-probabilities so as to capture the models internal representation of the distribution over token sequences. The representation is independent of the models computational structure. This representation yields the conditional probability of the response given the prompt, and of the response given the prompt with a token marginalized away. Our attribution score is the log of the ratio of these probabilities. We further compute the entropies of a single prompts token distributions, conditioned on the remaining context. The interplay between entropy and attribution score sheds light on LLM behavior. We evaluate 8 models across 7 prompts and investigate anomalies, token sensitivity, response stability, model stability, and training convergence, thereby improving interpretability and guiding users to focus on uncertain or unstable parts of the generation.
Shilpika Shilpika, Carlo Graziani, Bethany Lusch +2
Argonne Leadership Computing Facility Argonne National Laboratory Lemont, IL, USA · Mathematics and Computer Science Division Argonne National Laboratory Lemont, IL, USA · Department of Computer Science University of Illinois Chicago Chicago, IL, USA
Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality violations, where answers reflect errors in internal knowledge. Proper mitigation depends on knowing which source drives each answer. We study contributive attribution, i.e. the classification of the dominant knowledge source behind each output, and show that a simple linear probe trained on hidden representations can reliably identify it. We introduce AttriWiki, a self-supervised pipeline that automatically generates labelled training data by prompting models to recall withheld entities from memory or read them from context without relying on knowledge conflicts. Probes trained on AttriWiki achieve up to 0.96 Macro-F1 on Llama-3.1-8B, Mistral-7B, and Qwen-7B, transfer to SQuAD and WebQuestions with 0.94-0.99 Macro-F1, and generalise zero-shot to Tighidet et al. (2024)'s benchmark, outperforming their probe on conflicting settings without retraining. Furthermore, attribution mismatches raise error rates by up to 70%, though correct attribution does not guarantee correct answers, pointing to the need for broader detection frameworks.
Ivo Brink, Alexander Boer, Dennis Ulmer
KPMG NL / University of Amsterdam Amsterdam, The Netherlands · KPMG NL Amsterdam, The Netherlands · University of Amsterdam Amsterdam, The Netherlands