Distributional Validity of a Korean Synthetic Persona Panel: Evidence From the Korea Media Panel Survey
Authors: Howard Kim, Keuntae Cho
Organizations: College of AI Convergence, Seoul Cyber University, Seoul 01133, South Korea · Department of Management of Technology, Graduate School, Sungkyunkwan University, Suwon 16419, South Korea · Department of Systems Management Engineering, Sungkyunkwan University, Suwon 16419, South Korea
Large language model (LLM) personas are proposed as survey respondents, yet validation outside English-speaking contexts is scarce. We evaluate how well a Korean synthetic persona panel used to condition Gemini 3.5 Flash and EXAONE reproduces digital and artificial intelligence (AI) service-use distributions of the Korea Media Panel Survey. About 8,000 personas per model answered eight service-use items and eight attitudinal constructs; responses were compared with weighted survey estimates. The overall mean absolute error (MAE) was 14-19 percentage points (pp), with binary item-mean correlations of 0.70-0.91 across waves. Segment error across five axes was 14-18 pp, with between-group signed-error ranges of 49.6/34.7 pp (Gemini/EXAONE; 39.5/31.2 without the non-comparable teen cells). Errors were model-specific: an age stereotype (Gemini) versus an acquiescence-consistent level bias (EXAONE). Generative-AI overestimation was consistent with temporal misalignment; short-form underestimation was framing-sensitive and persisted under randomized order (both shown for Gemini). Post-hoc holdout calibration on 30% of the real data, with the correction form selected inside the calibration set, cut cell MAE from 18.3/15.1 to 4.9/4.4 pp, yet direct estimation from that subsample was more accurate than the calibrated panel (3.6 pp), a synthetic-informed shrinkage estimator beat its real-only counterpart by at most 0.7 pp, and the correction did not transfer competitively across waves. The calibrated panel kept an advantage only below roughly 250-860 real responses (at most 2.3 pp over a real-only shrinkage estimator) or, for one model, on unobserved segments. In this setting, synthetic panels are not survey substitutes; their value is diagnostic.
Figures & tables
Primary
Comparison
Model identifier
gemini-3.5-flash
LGAI-EXAONE/K-EXAONE-236B-A23B
Developer
Google
LG AI Research
Weights
closed
open (pinnable)
Access
Gemini Batch API
FriendliAI serverless, OpenAI-compatible
Temperature
1.0 (main) / 0.7
1.0 (main) / 0.7
top_p
1.0
1.0
TABLE I: Model and serving identifiers. Both systems were served in July 2026 (the randomized-order regeneration of Sec. IV-K on 3 September 2026); request payloads (system instruction, item text, decoding parameters) and run logs are archived in the reproducibility package.
Model
Wave
Personas
1st-attempt fail
Excluded
Excl. (%)
Gemini
2024
8,168
66 (0.81%)
0
0.00
Gemini
2025
7,938
5 (0.06%)
0
0.00
EXAONE
2024
8,168
n/a
3
0.04
EXAONE
2025
7,938
n/a
144
1.81
TABLE II: Format-violation regeneration and exclusion. Gemini (batch): requests failing the first round were resubmitted in up to two further rounds. EXAONE (live calls): up to three attempts per persona; first-attempt failures were not logged individually. Personas = sampled personas per run. The Gemini temperature-0.7 and demographic-only 2024 runs had 51 (0.6%) and 106 (1.3%) first-round failures and no exclusions; the EXAONE temperature-0.7 and demographic-only 2024 runs excluded 73 (0.9%) and 19 (0.2%) personas. Exclusions by sex-by-age cell (0.2–7.5%, worst cell men in their 40s, 2025) are given in Sec. III-D and in the package sheet of cell-level failures.
MAE replaces total variation distance as the primary measure; five-axis decomposition replaces the registered mixed-effects regression; real-data baselines replace the registered chance reference; the registered item-type contrasts (binary vs. multi-category; general digital vs. AI) are replaced by the binary-indicator vs. construct contrast, and the three multi-category AI items are not analyzed; of the registered robustness summaries, mean absolute differences are reported and cross-setting correlations are not; the EXAONE 2024 run raised the client’s output-token limit mid-run and repeated the three-attempt cycle once more for the affected personas (Sec. III-D); the survey’s OTT skip logic for the YouTube and short-form items was applied to the synthetic responses after the registered analyses had first been run (Sec. III-D); the eight registered age categories are analyzed as seven decade bands, with respondents under 10 (weighted share 1.1%) excluded from cell-based analyses
Wording-controlled regeneration; 2025 matched-wave regeneration and reference-year swap; randomized-order regeneration
Post-hoc (not preregistered)
RQ3 calibration in its entirety (all correction forms, nested selection, learning curve, entire-cell and temporal holdouts); real-only estimators; naive baselines and ablation; variance and stability analyses
TABLE III: Status of the reported analyses.
Fig. 1: RQ1 agreement, 2024. (a) Use rates: weighted survey estimates vs. Gemini and EXAONE (YouTube and short-form among OTT users, Sec. III-B); design-based sampling standard errors are below 1 pp per indicator (95% half-widths at most about 1.8 pp; Sec. III-E), smaller than the bar resolution. (b) Construct means (5-point): survey vs. synthetic.
Ref.
Model
MAE (pp)
Cosine
KL
JS
rbin
rcon
2024
Gemini
17.0
0.949
0.189
0.029
0.811
0.788
2024
EXAONE
13.8
0.971
0.066
0.018
0.912
0.500
2025
Gemini
19.4
0.914
0.235
0.052
0.703
–
2025
EXAONE
15.2
0.953
0.114
0.030
0.787
–
TABLE IV: RQ1 overall agreement metrics. Ref.: reference-survey year. The 2024 rows use the 2024-item panel; the 2025 rows use the matched-wave (2025-item) regeneration. Correlations are reported separately for the 8 binary items ( rbin ) and the 8 constructs ( rcon ; not measured in 2025), because pooling the two scales inflates the correlation. MAE 95% CIs (design-based household-cluster bootstrap, B=600 ): Gemini 2024 [16.5, 17.4], EXAONE 2024 [13.3, 14.4].
Axis
Gemini
EXAONE
Age
18.3
15.1
Sex
15.4
14.0
Education
17.3
14.7
Employment
15.2
13.9
Region (17 divisions)
15.4
15.6
Sex-by-age cells (combined)
18.3
15.1
TABLE V: RQ2 five-axis segment MAE (pp), combined sex-by-age cell MAE, the between-group signed-error range ( Re ), and its absolute-error complements. The five-axis rows are computed over each axis’s own groups; the lower block is over the 14 sex-by-age cells. All values are averaged across the eight indicators.
Fig. 2: Age structure of the error on the matched-wave 2025 panel. (a) Age slope of the signed error per indicator (seven pooled age bands; negative = older groups underestimated). (b) Short-form usage by age: gently declining survey trend vs. steep synthetic decline.
Model
Metric
Original
Neutral
Survey
Gemini
Short-form use
31.6
76.8
69.6
Gemini
P(short-form ∣ YouTube)
32.7
79.6
73.1
EXAONE
Short-form use
52.9
81.3
69.6
EXAONE
P(short-form ∣ YouTube)
59.9
92.6
73.1
TABLE VI: Wording-controlled regeneration (1,144 sampled personas per model; valid control/treatment responses 1,113/1,112 for Gemini and 969/998 for EXAONE; the same personas answered both arms, analyzed as independent samples; one call per persona without regeneration; all else fixed). Changing only the short-form item’s wording (removing “OTT” and naming the venues) moves short-form use by 45.2 pp (Gemini) and 28.3 pp (EXAONE) between the two arms. Survey rates are 2024 weighted estimates among OTT users; the synthetic arms are unconditional (Sec. IV-D), so the survey column is a reference level, not a like-for-like comparison. 95% CIs on the change exclude zero.
Fig. 3: Temporal comparisons. (a) Reference-year comparison: survey 2024/2025 vs. the 2024-item synthetic panels. (b) Held-out sex-by-age cell MAE for the uncorrected panel, the post-hoc linear age correction, and the nested selection: within-2024 (contemporaneous) vs. out-of-time (learned on 2024, tested on 2025). Dashed/dotted lines mark the target-year (2025) grand-mean (12.0) and prior-wave (6.7) references. Calibration cuts the error within-wave, but out-of-time the calibrated panel stays at or above both references (Sec. IV-G).
Model
Uncorrected
Global
Linear age
Nested
Gemini
18.3
14.2
8.3
4.9
EXAONE
15.1
8.4
6.7
4.4
Gemini (temporal)
20.2
17.8
13.9
12.0
EXAONE (temporal)
17.3
14.7
14.3
13.1
Grand-mean (2025)
12.0
Prior-wave (2024)
6.7
TABLE VII: RQ3 held-out sex-by-age cell MAE (pp) before and after calibration: global shift, the post-hoc linear age term, and the nested selection (form chosen inside the calibration set). The uncorrected column gives the full-sample cell MAE of Table V ; on the held-out cells of the same splits it is 18.3/15.2 pp. Top: 200 stratified within-2024 splits. Middle: temporal holdout (correction learned on 2024, evaluated on the 2025 panel against the 2025 reference estimates). Bottom: out-of-time references on the 2025 cells, the target-wave grand mean (oracle-style) and the prior wave. Both match or beat the temporally calibrated panel, so the correction reduces raw error across waves without transferring competitively.
Lin.
Nest.
Real-only
Win %
Fraction
Gem.
EXA.
Gem.
EXA.
Direct
Reg.
EB
GM
G/E
1% ( n≈80 )
10.2
8.6
9.3
8.5
13.4
10.1
10.7
12.7
76/94
2% ( n≈170 )
9.5
7.8
7.7
7.4
10.5
8.8
8.1
12.2
88/97
3% ( n≈250 )
9.1
7.4
6.8
6.5
8.8
8.2
6.9
12.0
96/99
5% ( n≈430 )
8.8
7.1
6.1
5.8
7.3
7.8
5.9
11.9
96/98
10% ( n≈860 )
8.5
6.8
5.5
5.1
5.5
7.4
4.9
11.8
50/81
TABLE VIII: Calibrated synthetic panel vs. real-only estimators on identical stratified splits (held-out sex-by-age cell MAE, pp; 200 splits per fraction; n = calibration respondents). Lin. = linear age form; Nest. = nested selection; Direct = weighted cell rates from the calibration set (empty cells fall back to the calibration set’s grand mean); Reg. = real-only regression (additive in age band and sex); EB = real-only empirical-Bayes shrinkage of the calibration cell means toward that regression line (no synthetic data, hence one value per fraction); GM = in-design calibration grand mean; Win = share of splits in which the nested-calibrated panel beat the better of Direct and Reg. (the share of splits in which each of the four estimators ranks first is reported in Sec. IV-G). Bottom row: entire-cell holdout (extrapolation to unobserved segments; direct estimation is undefined there).
Condition
Gemini
EXAONE
Synthetic (full persona)
18.3
15.1
Synthetic (demographic-only)
22.3
22.1
Grand-mean baseline
11.6
Prior-wave (2023) baseline (6 common items)
3.7
Calibrated synthetic (linear age, post-hoc)
8.3
6.7
Calibrated synthetic (nested selection)
4.9
4.4
TABLE IX: Sex-by-age cell MAE (pp): synthetic vs. baselines and ablation (2024, all eight indicators, identical 14 comparison cells throughout); the uncorrected, demographic-only, and baseline rows are full-sample values on the 14 cells, whereas the calibrated rows are held-out (200-split means). The prior-wave baseline covers only the six indicators fielded in 2023; on that common set, the values are uncorrected synthetic 13.3/13.7, grand mean 11.1, prior wave 3.7, and calibrated (linear) 6.9/6.8.
Indicator
Survey
F1
F2
R1
Δ [95% CI]
Holm p
Changed
AI use
13.8
41.1
40.5
46.9
+5.8 [3.9, 7.8]
0.008
11.1/5.3
OTT
89.2
64.1
63.5
62.7
−1.5 [ −3.4 , 0.5]
0.33
10.3/6.2
YouTube
93.2
100.0
100.0
100.0
0.0 [0.0, 0.0]
1.00
0.0/0.0
Short-form
69.8
33.4
31.9
60.1
+26.7 [21.6, 31.0]
0.008
32.1/13.6
SNS
61.4
49.5
46.7
53.2
+3.7 [1.6, 5.8]
0.012
10.5/7.1
Messenger
92.5
98.2
98.2
97.5
−0.7 [ −1.5 , 0.1]
0.23
2.0/1.3
TABLE X: Randomized-order regeneration (Gemini; 1,141 personas valid in all arms; post-stratified rates, %). F1/F2 = fixed codebook order, first and second pass (a third pass, F3, lies within 1.6 pp of both on every indicator; package sheet); R1 = random order per persona. The change column is R1 minus F1, computed from unrounded rates, with a persona-level bootstrap 95% CI and Holm-adjusted p (eight indicators). Changed = share of personas whose answer differed between F1 and R1 / between F1 and F2. Survey = 2024 weighted rate (ages 10 and over; YouTube and short-form among OTT users).
Figure 19Figure 20
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Metric (pp)
Gemini
EXAONE
RQ1 MAE, 2024
17.0 / 17.7
13.8 / 13.7
RQ1 MAE, 2025
19.4 / 20.4
15.2 / 15.2
RQ2 sex-by-age cell MAE
18.3 / 17.6
15.1 / 15.0
Re
49.6 / 39.5
34.7 / 31.2
Five-axis MAE range
15.2–18.3 / 15.6–17.5
13.9–15.6 / 13.6–16.1
Calibrated (linear age, post-hoc)
8.3 / 7.6
6.7 / 5.7
Appendix
TABLE XI: Teen-cell-excluded sensitivity: full sample / teen cells excluded (12 cells). Orderings are preserved except the Gemini entire-cell holdout (6.2 vs. 8.6 pp on 14 cells; 6.7 vs. 6.2 pp on the 12 non-teen cells); Re shrinks while headline error levels do not.
Indicator
10s
20s
30s
40s
50s
60s
70s+
AI use
+71
+47
+37
+29
+13
+4
+2
OTT
+2
+1
−3
−11
−40
−64
−40
YouTube
+9
+8
+6
+9
+6
+5
+3
Short-form
+11
−15
−51
−61
−60
−47
−20
SNS
+34
+6
−9
−31
−35
−26
−4
Messenger
+5
+1
+1
+1
+1
+6
+23
Appendix
TABLE XII: Gemini: signed error by age band (pp; synthetic − survey, 2024; YouTube and short-form among OTT users).
Indicator
10s
20s
30s
40s
50s
60s
70s+
AI use
+45
+28
+27
+28
+19
+13
+10
OTT
−4
−4
−8
−11
−12
−15
+21
YouTube
+8
+7
+5
+7
+2
−2
−7
Short-form
−9
−16
−23
−19
−13
−7
+14
SNS
+29
+3
−3
−2
+8
+21
+42
Messenger
−8
−10
−17
−24
−33
−41
−16
Appendix
TABLE XIII: EXAONE: signed error by age band (pp; synthetic − survey, 2024; YouTube and short-form among OTT users).
Model
Indicator
β0 [2.5, 97.5]
β1 [2.5, 97.5]
Resid. SD
HW
Gemini
AI use
62.9 [58.7, 67.0]
−11.3 [ −12.2 , −10.4 ]
5.9
2.3
Gemini
OTT
9.6 [7.7, 12.1]
−10.6 [ −11.4 , −9.9 ]
12.1
1.7
Gemini
YouTube
9.4 [5.9, 13.0]
−0.9 [ −1.7 , −0.2 ]
3.1
2.1
Gemini
Short-form
−16.8 [ −21.5 , −11.3 ]
−6.0 [ −7.4 , −4.7 ]
23.1
3.5
Gemini
SNS
12.4 [6.9, 18.1]
−7.3 [ −8.7 , −6.0 ]
19.7
3.3
Gemini
Messenger
−1.3 [ −3.5 , 1.3]
2.2 [1.4, 2.9]
5.7
1.9
Appendix
TABLE XIV: Linear age correction (Sec. III-F): fitted intercept β0 and slope β1 (pp; pp per decade band) per indicator, mean and 2.5/97.5 percentiles over the 200 calibration splits; residual SD of the full-data fit over the 14 cells; and the split-to-split half-width of the corrected cell estimate (HW, pp, averaged over cells).
Model
Mean ∣Δr∣
Max ∣Δr∣
r (vectors)
Mean ∣r∣ syn./survey
SD ratio (range)
Mean MAE
Gemini
0.085
0.252
0.01
0.70 / 0.68
0.71 (0.48–0.98)
0.37
EXAONE
0.108
0.320
0.35
0.62 / 0.68
0.85 (0.65–0.96)
0.39
Appendix
TABLE XV: Construct correlation structure (eight 5-point constructs, 2024): mean and maximum absolute difference between the synthetic and survey inter-construct correlations (28 pairs; survey correlations weighted), correlation between the two correlation vectors, mean absolute correlation, ratio of construct SDs (synthetic/survey), and MAE of construct means (1–5 scale). The construct MAE in this table compares unweighted synthetic construct means with the weighted survey means on the age-10-and-over sample (0.37/0.39); the value in Sec. IV-A (0.373/0.380) post-stratifies the synthetic means and uses the full survey sample, so the two are not the same statistic.
Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and pre-2023 survey histories, then testing them against the same respondents' held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.
Mumin Jia, Yilin Chen, Divya Sharma +1
Department of Mathematics and Statistics York University Toronto, Ontario M3J 1P3 · Department of Biostatistics University Health Network Toronto, ON M5G 2C4
Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent. Using real survey microdata, we show that this paradigm has a basic failure mode, and we set a distribution-first corrective against it, all measured with a deterministic, construct-validated verifier on non-WEIRD (Turkey-first) data. First, N independent LLM agents grounded on 2,414 real World Values Survey respondents fail to reproduce the population's response distribution: they pile onto a modal default (four scenarios x five seeds: concentration 0.36->0.69, entropy 1.46->0.77, 85% collapse, TVD=0.44), and the collapse is a predictable function of scenario structure (r=0.55 with a single-answer structure). Second, Verbalized Sampling (VS) fixes the field's chronic under-dispersion without training in three model families (fidelity +7 to +10; significant on Qwen, p=0.002, d=6.2), yet the same move universally overshoots into over-dispersion (SD-ratio 0.4-0.56 -> 1.26-1.37), a structural property of VS. Third, survey fidelity transfers only weakly to agentic behavior: in a single-model, single-domain booking task, a persona is dominated by a cheapest-default (~80%) that income modulates but does not override (comfort choice 0%->7%->32% across income bands). Fourth, a placebo-controlled memorization attack and an election backtest show VS keeps aggregate strength while subgroup and individual claims are contaminated by recall and underdetermination. We close with the corrective: model the distribution once (VS) and assign it to grounded characters at O(1) cost, with a budget-aware router whose honest AUC is 0.805, not the tautological 1.0 of a code-derived oracle. The central contribution needs no realism claim: it measures the internal inconsistency of the independent-agent route and the conditions under which the distribution-first route calibrates.
Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble human data. We study a fixed panel of sixteen lightweight persona-conditioned GPT-4.1 configurations in repeated strategic games. The panel met preregistered broad-reference condition-mean criteria in three of four repeated-game cells; the sole miss was 0.011 below the lower reference bound. Variation was strongly prompt-indexed, but its share depended on uncertainty assumptions: fixed-panel symmetric-Dirichlet sensitivities produced median between-prompt shares of 63%-71% under Jeffreys alpha=0.5 and 47%-53% under alpha=1, while finite-opportunity plug-in estimates were 85%-96%. Aggregate continuation-probability contrasts were +0.083 and +0.078, with conservative simultaneous 95% intervals [-0.171, +0.330] and [-0.181, +0.330]. The treatment jointly changed the continuation process and its textual representation. A separate wording-and-position operation shifted cooperation from 0/40 to 37/40 in the bare configuration, and a label conflict also revealed representation control. The original persona-level p13 result was not prospectively family-controlled, while a post-adjudication exact gate was structurally underpowered; p13 is therefore a replication target rather than a finding. External review exposed family-error, dependence, construct, and boundary-uncertainty defects, and zero-call reanalysis changed the interpretation without rewriting the historical record. The registered marginal criteria could be passed without precisely estimating the treatment-response object. A public capsule verifies 4,916 confirmatory Phase 3-5 runs with no live model calls. The results concern one fixed model-prompt panel and do not establish human substitutability.