Decision-only language models return a probability for every answer option instead of generating text, which makes them attractive as survey respondents and as judges. We audit the cultural values of one such model, TypeSafe's JEV, with the Values Survey Module 2013. We asked it the 24 items as 12 matched Saudi and 12 matched American personas and without a persona, in English and Arabic, under eight ways of formulating the request (288,000 answers). JEV's answers were highly repeatable (ICC 0.997), and without a persona they resembled those of its own American personas. When the persona was Saudi rather than American, the answers moved in the direction of the human Saudi-US difference, reproducing 87% of its size in English but 62% in Arabic, with long-term orientation reversed. A language cross shows that the smaller difference in Arabic comes from the language of the items, not from the language of the persona description. Age shifted the profiles about as much as nationality, gender shifted them more for Saudi than for American personas, and JEV was less confident in Arabic and for Saudi personas. These patterns held in every request design, although the model never generates text.
Figures & tables
Set-up
Request
1
Documented (headline)
Persona as a field of the state; item wording and a note; Choice
2
LLM-style prompt
System prompt as state; user prompt as question; Choice
3
No format sentence
As 2, without “IMPORTANT: Respond with ONLY …”
4
Score question
As 2, options as ordered levels
5
Score, reversed
As 4, levels in reverse order
6
Third person
“Description of the survey respondent: [persona].”; “Which answer would they choose?”
Table 1: The eight request set-ups. Each covers the 1,200 prompts, each sent 30 times.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Hofstede
Alm.
Saudi − US
Dimension
SA
US
SA
Hof.
Alm.
Power distance
95
40
72
+ 55
+ 32
Individualism
25
91
48
− 66
− 43
Masculinity
60
62
43
− 2
− 19
Uncertainty avoidance
80
46
64
+ 34
+ 18
Long-term orientation
36
26
27
+ 10
+ 1
Appendix
Table 2: Reference scores: Hofstede’s Saudi Arabia (SA) and United States (US) profiles and the Saudi profile of Almutairi et al. (2021) (Alm.), with the human Saudi-minus-US differences.
Arabic share
Cosine
PDI gap
Set-up
first
corr.
first
corr.
first
corr.
Documented
42.7%
62.0%
0.61
0.80
− 10.5
+ 26.0
LLM-style
51.3%
70.2%
0.57
0.72
− 7.4
+ 23.4
Appendix
Table 5: Arabic results with the answer options of the first runs and after the correction. PDI gap: Saudi minus American persona on power distance.
Set-up
ICC
Same top option (30 calls)
SD across calls
Mean confidence
Normalised entropy
Documented (headline)
0.997
85.0%
0.031
0.49
0.61
LLM-style prompt
0.997
88.7%
0.030
0.54
0.55
No format sentence
0.998
89.3%
0.026
0.58
0.49
Score question
0.996
85.4%
0.032
0.55
0.66
Score, levels reversed
0.995
83.3%
0.032
0.52
0.69
Third person
0.997
87.4%
0.030
0.58
0.49
Appendix
Table 6: Test–retest reliability over the 30 calls of each prompt and the shape of JEV’s answer distributions, by set-up (1,200 prompts each). *Framing in the other language than the items.
Share of the human difference
Default position
Confidence gap
Set-up
English
Arabic
English
Arabic
English − Arabic
American − Saudi
Documented (headline)
87.1%
62.0%
0.08
0.06
+ 0.090
+ 0.024
LLM-style prompt
103.5%
70.2%
− 0.25
0.02
+ 0.023
+ 0.020
No format sentence
101.3%
65.2%
− 0.17
− 0.08
+ 0.008
+ 0.018
Score question
106.8%
69.4%
− 0.04
− 0.00
+ 0.063
+ 0.035
Score, levels reversed
99.8%
72.2%
− 0.05
− 0.01
+ 0.090
+ 0.040
Appendix
Table 7: Main findings by set-up. Default position: 0 = JEV’s American personas, 1 = its Saudi personas. Confidence gaps: mean differences over matched prompts. *English and Arabic refer to the items; the framing is in the other language.
Personas, English
Personas, Arabic
No persona, Arabic − English
Set-up
In order
Out of order
In order
Out of order
In order
Out of order
Documented (headline)
4 (0.33)
MAS, LTO
4 (0.33)
MAS, LTO
1 ( − 0.67)
PDI, MAS, UAI, LTO, IVR
LLM-style prompt
4 (0.33)
MAS, LTO
4 (0.33)
MAS, LTO
3 (0.00)
MAS, LTO, IVR
No format sentence
5 (0.67)
LTO
4 (0.33)
MAS, LTO
3 (0.00)
MAS, LTO, IVR
Score question
5 (0.67)
LTO
4 (0.33)
MAS, LTO
4 (0.33)
LTO, IVR
Score, levels reversed
4 (0.33)
MAS, LTO
4 (0.33)
MAS, LTO
4 (0.33)
LTO, IVR
Appendix
Table 8: Rank test of Masoud et al. (2025) by set-up: number of dimensions (of six) on which JEV orders Saudi Arabia and the United States as Hofstede’s scores do, with the mean Kendall’s τ in parentheses, and the dimensions out of order. Personas: Saudi minus American personas (means over the 12 matched profiles). No persona: the Arabic minus the English no-persona profile, with the language standing for the country. The Saudi profile of Almutairi et al. (2021) gives the same orders. *English and Arabic refer to the items; the framing is in the other language.
Dimension
Hofstede
English
Arabic
Power distance
+ 55
+ 44.9 [36.9, 53.0]
+ 23.4 [20.4, 26.4]
Individualism
− 66
− 80.0 [ − 90.3, − 69.7]
− 57.9 [ − 66.9, − 48.9]
Masculinity
− 2
+ 1.4 [ − 4.3, 7.2]
+ 8.4 [3.1, 13.7]
Uncertainty avoidance
+ 34
+ 23.7 [14.7, 32.7]
+ 8.4 [6.9, 10.0]
Long-term orientation
+ 10
− 31.7 [ − 38.8, − 24.6]
− 16.7 [ − 22.6, − 10.9]
Indulgence
− 16
− 60.6 [ − 69.5, − 51.7]
− 64.4 [ − 74.2, − 54.5]
Appendix
Table 9: Saudi minus American persona by dimension with the LLM-style prompt (means over 12 matched profiles, 95% CIs), with the Hofstede difference.
Saudi − American answer
Confidence
Item
Dim.
Content
English
Arabic
English
Arabic
1
IDV
time for personal or home life
+ 0.01
− 0.06
0.59
0.42
2
PDI
a boss you can respect
− 0.22
− 0.51
0.59
0.35
3
MAS
recognition for good performance
− 0.09
− 0.11
0.43
0.41
4
IDV
security of employment
− 0.44
− 0.34
0.55
0.44
5
MAS
pleasant people to work with
− 0.07
− 0.12
0.68
0.41
Appendix
Table 10: Item-level results in the documented set-up: mean difference between Saudi and American personas in the expected answer (1–5 scale; negative = Saudi personas chose options nearer the first, e.g. more important or prouder), and JEV’s mean confidence, by language.
When you ask an AI assistant for advice about your career, your marriage, or a conflict with your family, does it give you the same answer regardless of where you are from? We tested this systematically by presenting three leading AI systems (Claude Sonnet 4.5, GPT-5.4, and Gemini 2.5 Flash) with ten real-life personal dilemmas, framed for users from 10 countries across 5 continents in 7 languages (n=840 scored responses). We compared AI advice against World Values Survey Wave 7 data measuring what people in each country actually believe. All three AI systems consistently gave Western-style, individualist advice even to users from societies that prioritize family, community, and authority, significantly more so than local values would predict (mean gap +0.76 on a 1-5 scale; t=15.65, p<0.001). The gap is largest for Nigeria (+1.85) and India (+0.82). Japan is the sole exception: AI systems treated Japanese users as more group-oriented than surveys show, revealing that AI encodes outdated stereotypes. Claude and GPT-5.4 show nearly identical bias magnitude, while Gemini is lower but still significant. The models diverge in mechanism: Claude shifts further collectivist in the user's native language; Gemini shifts more individualist; GPT-5.4 responds only to stated country identity. These findings point to a systemic homogenization of values across frontier AI. Data, code, and scoring pipeline are openly released.
Large language models (LLMs) are increasingly used to simulate human opinions and survey responses, but their ability to reproduce population responses across cultures remains limited. Existing persona-based prompting methods typically rely on sociodemographic or personality traits, which are only indirect proxies for the values that shape human responses. We propose a value-based persona construction method that derives textual descriptors from survey responses capturing core cultural dimensions. By sampling value profiles from target populations and aggregating LLM responses across personas, we obtain population-level predictions grounded in observed value distributions. We further introduce a calibration procedure that improves response diversity while preserving estimated opinions. We show that our approach reduces prediction error across countries, with the largest improvements observed in underrepresented populations. This substantially narrows the performance gap between countries aligned with dominant LLM priors and those that are less represented in training data, while also yielding response distributions that closely match human diversity.
Axel Abels, Elias Fernandez Domingos, Apurva Shah +1
1Machine Learning Group, Universit´e Libre de Bruxelles · 2AI Lab, Vrije Universiteit Brussel · 3FARI Institute, Universit´e Libre de Bruxelles - Vrije Universiteit Brussel +2
The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity to reinforce harmful societal stereotypes. While significant attention has been paid to such fairness concerns in the context of social biases, relatively little prior work has examined the presence of stereotypes in LVLMs related to cultural contexts such as religion, nationality, and socioeconomic status. In this work, we aim to narrow this gap by investigating how cultural contexts depicted in images influence the judgments LVLMs make about a person's moral, ethical, and political values. We conduct a multi-dimensional analysis of such value judgments in nine LVLMs using counterfactual image sets, which depict the same person across different cultural contexts. Our evaluation framework pairs descriptive analyses (Moral Foundations Theory categorization, lexical analyses, and value sensitivity) with a novel grounding analysis that compares LVLM cross-context variation against two large-scale human surveys (MFQ-2 and WVS Wave 7). Across 4.8 million LVLM generations, we identify three bias patterns that replicate across architecturally diverse models: an inversion of the socioeconomic-status-to-Authority relationship found in WVS, and two race-conditional failures that override cultural context cues when depicting Middle Eastern persons. Additional ablations show that the socioeconomic-status-to-Authority inversion bias is amplified by image conditioning and persists across different model sizes.