Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity within countries but also whether differences across countries are preserved. However, existing distance-based metrics such as Jensen--Shannon divergence (JSD) do not directly capture such cross-country differences. To address this limitation, we introduce Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration. CDP identifies reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature. To evaluate CDP, we conduct experiments across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test. The results reveal a systematic discrepancy between conventional fidelity metrics and CDP. Controlled experiments show that CDP changes monotonically as cross-country divergence is attenuated or amplified, while the corresponding changes in JSD remain relatively small. In our audit of real LLM generations, DeepPersona-Inspired prompting is frequently favored by conventional fidelity metrics but exhibits the strongest flattening in every model--domain block. CDP thus complements fidelity metrics by directly quantifying the attenuation or amplification of cross-country divergence.
Figures & tables
Figure 1: Motivations of CDP.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
ID
Construct and survey item
Response scale
Q45
Respect for Authority If greater respect for authority takes place in the near future, do you think it would be a good thing, a bad thing, or you don’t mind?
1 = A good thing; 2 = Don’t mind; 3 = A bad thing
Q46
Feeling of Happiness Taking all things together, rate how happy you would say you are.
1 = Very happy; 2 = Quite happy; 3 = Not very happy; 4 = Not at all happy
Q57
Trust on People Generally speaking, would you say that most people can be trusted or that you need to be very careful in dealing with people?
1 = Most people can be trusted; 2 = Need to be very careful
Q184
Justifiability of Abortion How justifiable do you think abortion is?
1 = Never justifiable … 10 = Always justifiable
Q218
Petition Signing Have you signed a petition?
1 = Have done; 2 = Might do; 3 = Would never do
Q254
Pride of Nationality How proud are you to be your nationality?
1 = Very proud; 2 = Quite proud; 3 = Not very proud; 4 = Not at all proud
Appendix
Table 2: World Values Survey items.
ID
Item
ID
Item
EXT1
I am the life of the party.
AGR6
I have a soft heart.
EXT2
I don’t talk a lot.
AGR7
I am not really interested in others.
EXT3
I feel comfortable around people.
AGR8
I take time out for others.
EXT4
I keep in the background.
AGR9
I feel others’ emotions.
EXT5
I start conversations.
AGR10
I make people feel at ease.
EXT6
I have little to say.
CSN1
I am always prepared.
Appendix
Table 3: Big Five personality items rated on a 5-point scale (1 = Disagree strongly; 5 = Agree strongly).
WVS
Big Five
Argentina
Australia
Germany
India
Kenya
United States
Argentina
Australia
India
1,003
1,813
1,528
1,692
1,266
2,596
486
8,584
2,820
Appendix
Table 4: Country-level human sample sizes for WVS and Big Five.
CDPc cutoff t
Conditions
Methods represented
0.4
31
DeepP-I, PHub-I
0.5
37
DeepP-I, PHub-I
0.6
42
DeepP-I, PHub-I
Appendix
Table 9: Country-level discordance across CDPc cutoffs. Conditions are counted among the same 108 country–model–method cells. PHub-I: PersonaHub-Inspired; DeepP-I: DeepPersona-Inspired.