Distributional Validity of a Korean Synthetic Persona Panel: Evidence From the Korea Media Panel Survey
Authors: Howard Kim, Keuntae Cho
Organizations: College of AI Convergence, Seoul Cyber University, Seoul 01133, South Korea · Department of Management of Technology, Graduate School, Sungkyunkwan University, Suwon 16419, South Korea · Department of Systems Management Engineering, Sungkyunkwan University, Suwon 16419, South Korea
Large language model (LLM) personas are proposed as survey respondents, yet validation outside English-speaking contexts is scarce. We evaluate how well a Korean synthetic persona panel used to condition Gemini 3.5 Flash and EXAONE reproduces digital and artificial intelligence (AI) service-use distributions of the Korea Media Panel Survey. About 8,000 personas per model answered eight service-use items and eight attitudinal constructs; responses were compared with weighted survey estimates. The overall mean absolute error (MAE) was 14-19 percentage points (pp), with binary item-mean correlations of 0.70-0.91 across waves. Segment error across five axes was 14-18 pp, with between-group signed-error ranges of 49.6/34.7 pp (Gemini/EXAONE; 39.5/31.2 without the non-comparable teen cells). Errors were model-specific: an age stereotype (Gemini) versus an acquiescence-consistent level bias (EXAONE). Generative-AI overestimation was consistent with temporal misalignment; short-form underestimation was framing-sensitive and persisted under randomized order (both shown for Gemini). Post-hoc holdout calibration on 30% of the real data, with the correction form selected inside the calibration set, cut cell MAE from 18.3/15.1 to 4.9/4.4 pp, yet direct estimation from that subsample was more accurate than the calibrated panel (3.6 pp), a synthetic-informed shrinkage estimator beat its real-only counterpart by at most 0.7 pp, and the correction did not transfer competitively across waves. The calibrated panel kept an advantage only below roughly 250-860 real responses (at most 2.3 pp over a real-only shrinkage estimator) or, for one model, on unobserved segments. In this setting, synthetic panels are not survey substitutes; their value is diagnostic.
Figures & tables
Primary
Comparison
Model identifier
gemini-3.5-flash
LGAI-EXAONE/K-EXAONE-236B-A23B
Developer
Google
LG AI Research
Weights
closed
open (pinnable)
Access
Gemini Batch API
FriendliAI serverless, OpenAI-compatible
Temperature
1.0 (main) / 0.7
1.0 (main) / 0.7
top_p
1.0
1.0
TABLE I: Model and serving identifiers. Both systems were served in July 2026 (the randomized-order regeneration of Sec. IV-K on 3 September 2026); request payloads (system instruction, item text, decoding parameters) and run logs are archived in the reproducibility package.
Model
Wave
Personas
1st-attempt fail
Excluded
Excl. (%)
Gemini
2024
8,168
66 (0.81%)
0
0.00
Gemini
2025
7,938
5 (0.06%)
0
0.00
EXAONE
2024
8,168
n/a
3
0.04
EXAONE
2025
7,938
n/a
144
1.81
TABLE II: Format-violation regeneration and exclusion. Gemini (batch): requests failing the first round were resubmitted in up to two further rounds. EXAONE (live calls): up to three attempts per persona; first-attempt failures were not logged individually. Personas = sampled personas per run. The Gemini temperature-0.7 and demographic-only 2024 runs had 51 (0.6%) and 106 (1.3%) first-round failures and no exclusions; the EXAONE temperature-0.7 and demographic-only 2024 runs excluded 73 (0.9%) and 19 (0.2%) personas. Exclusions by sex-by-age cell (0.2–7.5%, worst cell men in their 40s, 2025) are given in Sec. III-D and in the package sheet of cell-level failures.
MAE replaces total variation distance as the primary measure; five-axis decomposition replaces the registered mixed-effects regression; real-data baselines replace the registered chance reference; the registered item-type contrasts (binary vs. multi-category; general digital vs. AI) are replaced by the binary-indicator vs. construct contrast, and the three multi-category AI items are not analyzed; of the registered robustness summaries, mean absolute differences are reported and cross-setting correlations are not; the EXAONE 2024 run raised the client’s output-token limit mid-run and repeated the three-attempt cycle once more for the affected personas (Sec. III-D); the survey’s OTT skip logic for the YouTube and short-form items was applied to the synthetic responses after the registered analyses had first been run (Sec. III-D); the eight registered age categories are analyzed as seven decade bands, with respondents under 10 (weighted share 1.1%) excluded from cell-based analyses
Wording-controlled regeneration; 2025 matched-wave regeneration and reference-year swap; randomized-order regeneration
Post-hoc (not preregistered)
RQ3 calibration in its entirety (all correction forms, nested selection, learning curve, entire-cell and temporal holdouts); real-only estimators; naive baselines and ablation; variance and stability analyses
TABLE III: Status of the reported analyses.
Fig. 1: RQ1 agreement, 2024. (a) Use rates: weighted survey estimates vs. Gemini and EXAONE (YouTube and short-form among OTT users, Sec. III-B); design-based sampling standard errors are below 1 pp per indicator (95% half-widths at most about 1.8 pp; Sec. III-E), smaller than the bar resolution. (b) Construct means (5-point): survey vs. synthetic.
Ref.
Model
MAE (pp)
Cosine
KL
JS
rbin
rcon
2024
Gemini
17.0
0.949
0.189
0.029
0.811
0.788
2024
EXAONE
13.8
0.971
0.066
0.018
0.912
0.500
2025
Gemini
19.4
0.914
0.235
0.052
0.703
–
2025
EXAONE
15.2
0.953
0.114
0.030
0.787
–
TABLE IV: RQ1 overall agreement metrics. Ref.: reference-survey year. The 2024 rows use the 2024-item panel; the 2025 rows use the matched-wave (2025-item) regeneration. Correlations are reported separately for the 8 binary items ( rbin ) and the 8 constructs ( rcon ; not measured in 2025), because pooling the two scales inflates the correlation. MAE 95% CIs (design-based household-cluster bootstrap, B=600 ): Gemini 2024 [16.5, 17.4], EXAONE 2024 [13.3, 14.4].
Axis
Gemini
EXAONE
Age
18.3
15.1
Sex
15.4
14.0
Education
17.3
14.7
Employment
15.2
13.9
Region (17 divisions)
15.4
15.6
Sex-by-age cells (combined)
18.3
15.1
TABLE V: RQ2 five-axis segment MAE (pp), combined sex-by-age cell MAE, the between-group signed-error range ( Re ), and its absolute-error complements. The five-axis rows are computed over each axis’s own groups; the lower block is over the 14 sex-by-age cells. All values are averaged across the eight indicators.
Fig. 2: Age structure of the error on the matched-wave 2025 panel. (a) Age slope of the signed error per indicator (seven pooled age bands; negative = older groups underestimated). (b) Short-form usage by age: gently declining survey trend vs. steep synthetic decline.
Model
Metric
Original
Neutral
Survey
Gemini
Short-form use
31.6
76.8
69.6
Gemini
P(short-form ∣ YouTube)
32.7
79.6
73.1
EXAONE
Short-form use
52.9
81.3
69.6
EXAONE
P(short-form ∣ YouTube)
59.9
92.6
73.1
TABLE VI: Wording-controlled regeneration (1,144 sampled personas per model; valid control/treatment responses 1,113/1,112 for Gemini and 969/998 for EXAONE; the same personas answered both arms, analyzed as independent samples; one call per persona without regeneration; all else fixed). Changing only the short-form item’s wording (removing “OTT” and naming the venues) moves short-form use by 45.2 pp (Gemini) and 28.3 pp (EXAONE) between the two arms. Survey rates are 2024 weighted estimates among OTT users; the synthetic arms are unconditional (Sec. IV-D), so the survey column is a reference level, not a like-for-like comparison. 95% CIs on the change exclude zero.
Fig. 3: Temporal comparisons. (a) Reference-year comparison: survey 2024/2025 vs. the 2024-item synthetic panels. (b) Held-out sex-by-age cell MAE for the uncorrected panel, the post-hoc linear age correction, and the nested selection: within-2024 (contemporaneous) vs. out-of-time (learned on 2024, tested on 2025). Dashed/dotted lines mark the target-year (2025) grand-mean (12.0) and prior-wave (6.7) references. Calibration cuts the error within-wave, but out-of-time the calibrated panel stays at or above both references (Sec. IV-G).
Model
Uncorrected
Global
Linear age
Nested
Gemini
18.3
14.2
8.3
4.9
EXAONE
15.1
8.4
6.7
4.4
Gemini (temporal)
20.2
17.8
13.9
12.0
EXAONE (temporal)
17.3
14.7
14.3
13.1
Grand-mean (2025)
12.0
Prior-wave (2024)
6.7
TABLE VII: RQ3 held-out sex-by-age cell MAE (pp) before and after calibration: global shift, the post-hoc linear age term, and the nested selection (form chosen inside the calibration set). The uncorrected column gives the full-sample cell MAE of Table V ; on the held-out cells of the same splits it is 18.3/15.2 pp. Top: 200 stratified within-2024 splits. Middle: temporal holdout (correction learned on 2024, evaluated on the 2025 panel against the 2025 reference estimates). Bottom: out-of-time references on the 2025 cells, the target-wave grand mean (oracle-style) and the prior wave. Both match or beat the temporally calibrated panel, so the correction reduces raw error across waves without transferring competitively.
Lin.
Nest.
Real-only
Win %
Fraction
Gem.
EXA.
Gem.
EXA.
Direct
Reg.
EB
GM
G/E
1% ( n≈80 )
10.2
8.6
9.3
8.5
13.4
10.1
10.7
12.7
76/94
2% ( n≈170 )
9.5
7.8
7.7
7.4
10.5
8.8
8.1
12.2
88/97
3% ( n≈250 )
9.1
7.4
6.8
6.5
8.8
8.2
6.9
12.0
96/99
5% ( n≈430 )
8.8
7.1
6.1
5.8
7.3
7.8
5.9
11.9
96/98
10% ( n≈860 )
8.5
6.8
5.5
5.1
5.5
7.4
4.9
11.8
50/81
TABLE VIII: Calibrated synthetic panel vs. real-only estimators on identical stratified splits (held-out sex-by-age cell MAE, pp; 200 splits per fraction; n = calibration respondents). Lin. = linear age form; Nest. = nested selection; Direct = weighted cell rates from the calibration set (empty cells fall back to the calibration set’s grand mean); Reg. = real-only regression (additive in age band and sex); EB = real-only empirical-Bayes shrinkage of the calibration cell means toward that regression line (no synthetic data, hence one value per fraction); GM = in-design calibration grand mean; Win = share of splits in which the nested-calibrated panel beat the better of Direct and Reg. (the share of splits in which each of the four estimators ranks first is reported in Sec. IV-G). Bottom row: entire-cell holdout (extrapolation to unobserved segments; direct estimation is undefined there).
Condition
Gemini
EXAONE
Synthetic (full persona)
18.3
15.1
Synthetic (demographic-only)
22.3
22.1
Grand-mean baseline
11.6
Prior-wave (2023) baseline (6 common items)
3.7
Calibrated synthetic (linear age, post-hoc)
8.3
6.7
Calibrated synthetic (nested selection)
4.9
4.4
TABLE IX: Sex-by-age cell MAE (pp): synthetic vs. baselines and ablation (2024, all eight indicators, identical 14 comparison cells throughout); the uncorrected, demographic-only, and baseline rows are full-sample values on the 14 cells, whereas the calibrated rows are held-out (200-split means). The prior-wave baseline covers only the six indicators fielded in 2023; on that common set, the values are uncorrected synthetic 13.3/13.7, grand mean 11.1, prior wave 3.7, and calibrated (linear) 6.9/6.8.
Indicator
Survey
F1
F2
R1
Δ [95% CI]
Holm p
Changed
AI use
13.8
41.1
40.5
46.9
+5.8 [3.9, 7.8]
0.008
11.1/5.3
OTT
89.2
64.1
63.5
62.7
−1.5 [ −3.4 , 0.5]
0.33
10.3/6.2
YouTube
93.2
100.0
100.0
100.0
0.0 [0.0, 0.0]
1.00
0.0/0.0
Short-form
69.8
33.4
31.9
60.1
+26.7 [21.6, 31.0]
0.008
32.1/13.6
SNS
61.4
49.5
46.7
53.2
+3.7 [1.6, 5.8]
0.012
10.5/7.1
Messenger
92.5
98.2
98.2
97.5
−0.7 [ −1.5 , 0.1]
0.23
2.0/1.3
TABLE X: Randomized-order regeneration (Gemini; 1,141 personas valid in all arms; post-stratified rates, %). F1/F2 = fixed codebook order, first and second pass (a third pass, F3, lies within 1.6 pp of both on every indicator; package sheet); R1 = random order per persona. The change column is R1 minus F1, computed from unrounded rates, with a persona-level bootstrap 95% CI and Holm-adjusted p (eight indicators). Changed = share of personas whose answer differed between F1 and R1 / between F1 and F2. Survey = 2024 weighted rate (ages 10 and over; YouTube and short-form among OTT users).
Figure 19Figure 20
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Metric (pp)
Gemini
EXAONE
RQ1 MAE, 2024
17.0 / 17.7
13.8 / 13.7
RQ1 MAE, 2025
19.4 / 20.4
15.2 / 15.2
RQ2 sex-by-age cell MAE
18.3 / 17.6
15.1 / 15.0
Re
49.6 / 39.5
34.7 / 31.2
Five-axis MAE range
15.2–18.3 / 15.6–17.5
13.9–15.6 / 13.6–16.1
Calibrated (linear age, post-hoc)
8.3 / 7.6
6.7 / 5.7
Appendix
TABLE XI: Teen-cell-excluded sensitivity: full sample / teen cells excluded (12 cells). Orderings are preserved except the Gemini entire-cell holdout (6.2 vs. 8.6 pp on 14 cells; 6.7 vs. 6.2 pp on the 12 non-teen cells); Re shrinks while headline error levels do not.
Indicator
10s
20s
30s
40s
50s
60s
70s+
AI use
+71
+47
+37
+29
+13
+4
+2
OTT
+2
+1
−3
−11
−40
−64
−40
YouTube
+9
+8
+6
+9
+6
+5
+3
Short-form
+11
−15
−51
−61
−60
−47
−20
SNS
+34
+6
−9
−31
−35
−26
−4
Messenger
+5
+1
+1
+1
+1
+6
+23
Appendix
TABLE XII: Gemini: signed error by age band (pp; synthetic − survey, 2024; YouTube and short-form among OTT users).
Indicator
10s
20s
30s
40s
50s
60s
70s+
AI use
+45
+28
+27
+28
+19
+13
+10
OTT
−4
−4
−8
−11
−12
−15
+21
YouTube
+8
+7
+5
+7
+2
−2
−7
Short-form
−9
−16
−23
−19
−13
−7
+14
SNS
+29
+3
−3
−2
+8
+21
+42
Messenger
−8
−10
−17
−24
−33
−41
−16
Appendix
TABLE XIII: EXAONE: signed error by age band (pp; synthetic − survey, 2024; YouTube and short-form among OTT users).
Model
Indicator
β0 [2.5, 97.5]
β1 [2.5, 97.5]
Resid. SD
HW
Gemini
AI use
62.9 [58.7, 67.0]
−11.3 [ −12.2 , −10.4 ]
5.9
2.3
Gemini
OTT
9.6 [7.7, 12.1]
−10.6 [ −11.4 , −9.9 ]
12.1
1.7
Gemini
YouTube
9.4 [5.9, 13.0]
−0.9 [ −1.7 , −0.2 ]
3.1
2.1
Gemini
Short-form
−16.8 [ −21.5 , −11.3 ]
−6.0 [ −7.4 , −4.7 ]
23.1
3.5
Gemini
SNS
12.4 [6.9, 18.1]
−7.3 [ −8.7 , −6.0 ]
19.7
3.3
Gemini
Messenger
−1.3 [ −3.5 , 1.3]
2.2 [1.4, 2.9]
5.7
1.9
Appendix
TABLE XIV: Linear age correction (Sec. III-F): fitted intercept β0 and slope β1 (pp; pp per decade band) per indicator, mean and 2.5/97.5 percentiles over the 200 calibration splits; residual SD of the full-data fit over the 14 cells; and the split-to-split half-width of the corrected cell estimate (HW, pp, averaged over cells).
Model
Mean ∣Δr∣
Max ∣Δr∣
r (vectors)
Mean ∣r∣ syn./survey
SD ratio (range)
Mean MAE
Gemini
0.085
0.252
0.01
0.70 / 0.68
0.71 (0.48–0.98)
0.37
EXAONE
0.108
0.320
0.35
0.62 / 0.68
0.85 (0.65–0.96)
0.39
Appendix
TABLE XV: Construct correlation structure (eight 5-point constructs, 2024): mean and maximum absolute difference between the synthetic and survey inter-construct correlations (28 pairs; survey correlations weighted), correlation between the two correlation vectors, mean absolute correlation, ratio of construct SDs (synthetic/survey), and MAE of construct means (1–5 scale). The construct MAE in this table compares unweighted synthetic construct means with the weighted survey means on the age-10-and-over sample (0.37/0.39); the value in Sec. IV-A (0.373/0.380) post-stratifies the synthetic means and uses the full survey sample, so the two are not the same statistic.
Department of Mathematics and Statistics York University Toronto, Ontario M3J 1P3 · Department of Biostatistics University Health Network Toronto, ON M5G 2C4