Authors: Fan Huang, Minsuk Kim, C. Tyler Diggans, Filippo Radicchi
Organizations: Indiana University Bloomington Bloomington, IN, USA · Massachusetts Institute of Technology Cambridge, MA, USA · Root Dynamix LLC Sewickley, PA, USA
Large Language Models (LLMs) are increasingly used as probabilistic generators for simulation, synthetic data generation, and decision support in settings where real-world data are unavailable. Yet, the structure and reliability of the distributions they produce remain understudied. Here, we systematically analyze LLM-generated distributions of preferences for air travel, restaurants, and consumer products. Encouragingly, all models considered in our analysis exhibit self-coherence, with the most probable outcomes stabilizing rapidly under repeated sampling. At the same time, we observe substantial discordance across both model families and scales, with little consensus even among their most probable outcomes. These patterns hold across nine open-weight models, three choice domains, and show robustness under temperature changes, greedy decoding, and perturbations of prompt and ordering. Our findings indicate that outcomes are influenced more by the choice of model than by the wording of the prompt, challenging the common assumption that sufficiently capable LLMs produce similar preference distributions when used as stand-ins for survey respondents.
Figures & tables
Context/Model
Unique@1
Unique@20
Stable@
IND
19
32
Iter 3
FWA
16
42
Iter 6
ATL
15
43
Iter 5
ORD
18
34
Iter 3
Model Size Comparison (IND)
Baseline (70B)
19
32
Iter 3
Table 1: Convergence results for travel demand (US airports). Unique@ R : cumulative support size, distinct destinations with non-zero probability pooled over the first R samples (Section 3.5 ). Stable@ : the first sample from which this cumulative support plateaus (changes by at most two destinations over three successive samples); a support-size stabilization signal, complementary to and distinct from the top- k stabilization point R∗ (Section 3.5 ).
Iter.
JStotal
JSmode
JStail
Δtail−mode
1–3
0.12±0.06
0.07±0.06
0.20±0.07
+0.13±0.03
3–5
0.19±0.09
0.13±0.08
0.31±0.11
+0.18±0.05
5–10
0.08±0.03
0.06±0.04
0.12±0.04
+0.06±0.03
10–20
0.09±0.03
0.06±0.03
0.13±0.04
+0.07±0.04
Table 2: JS divergence by convergence phase on the full distribution, the renormalized top-5 mode region, and its complement (US airport travel demand). Cells are mean ± half-CI (§ 4 ) over consecutive-sample comparisons pooled across the four contexts ( N=12,8,20,34 by phase). Δtail−mode , the paired difference, is the appropriate test of tail dominance since both regions share a sample pair: it excludes zero in every phase, whereas the marginal intervals overlap in three of the four.
Model Pair
JS Div.
Agr.@5
Top1
70B vs 405B
0.36±0.04
0.12±0.05
0.30±0.20
70B vs 8B
0.47±0.05
0.08±0.04
0.05±0.07
405B vs 8B
0.50±0.04
0.08±0.03
0.20±0.18
Table 3: Scale discordance metrics across three Llama sizes (US airport travel demand). Mean ± half-CI over 20 personas; see § 4 for the precision convention.
Figure 1: Per-sample distributional signatures across three LLM sizes (Section 4 resampling protocol; the CI band for the number of sampled distributions equal to 20 is by definition zero because there is only one sample equal to the full pool, so the informative part of each curve is the small- R region). (a) Unique destinations: 8B (mean 71.6) is 1.58 × the 405B (45.2) and 2.45 × the 70B (29.3). (b) Maximum destination probability: 70B and 405B concentrate 0.171 and 0.174, respectively; 8B is more diffuse (0.090). (c) Std of probabilities: 70B and 405B at 0.038; 8B tighter (0.019). The three models occupy structurally different operating points, not refinements of one another: across the 20 personas no two share an identical 70B/405B top-5 (mean top-5 overlap 0.12 ) and even the single dominant destination differs in most cases (top-1 match 0.30 ; Table 3 ), so the larger model re-weights which destinations dominate rather than sharpening the same modes.
Family
S-M
M-L
S-L
Llama
0.47±0.05
0.36±0.04
0.50±0.04
Qwen
0.14 †
0.15 †
0.23 †
Mistral
0.22 †
0.14 †
0.28 †
Table 4: Cross-family scale discordance (JS divergence; US airport travel demand). Llama: mean ± half-CI over 20 personas (same protocol as Table 3 ). † single-context point estimate (IND only); airport names are normalized to IATA codes before comparison (Appendix I ).
Perturbation Type
JS Div.
N
Persona (with vs without)
0.13±0.01
298
Age (within education)
0.33±0.05
24
Education (within age)
0.26±0.03
60
Context (regional vs hub)
0.31±0.04
40
Table 5: Prompt sensitivity effect sizes (JS divergence; US airport travel demand). Mean ± half-CI over N pairwise comparisons; see § 4 for the precision convention.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Method
T
Sup.
Entropy
Top-5
JS
Single-choice
0.2
1
0.00±0.00
1.00±0.00
0.64±0.02
Single-choice
1.0
2
0.18±0.10
1.00±0.00
0.61±0.04
Verbalized
0.2
35
3.97±0.29
0.58±0.06
(ref.)
Verbalized
1.0
34
4.25±0.20
0.46±0.04
(ref.)
Appendix
Table 6: Comparison of distribution elicitation methods for IND departures (Llama-3.3-70B). Sup. : unique destinations observed. Entropy : Shannon entropy in bits. Top-5 : probability mass in top 5 destinations. JS : Jensen–Shannon divergence against the verbalized baseline at the same temperature; “(ref.)” marks the verbalized rows, which are that baseline (self-divergence 0 ). Entropy, Top-5 and JS are mean ± half-CI over 10,000 bootstrap resamples of the underlying calls (20 verbalized calls at T=0.2 and 18 at T=1.0 ; 300 single-choice draws at each). Support is an exact count over the full sample and is given without an interval, because resampling calls can only remove distinct destinations and never add them, so its bootstrap is biased downward. The zero intervals on the single-choice rows are structural rather than rounding: at T=0.2 every draw returned the same destination, and with a support of one or two all mass necessarily falls inside the top five. Single-choice sampling produces degenerate distributions at both temperatures, while verbalized elicitation captures rich distributional structure in a single query.
Family
Model ID
Params
Role
Llama
Llama-3.1-8B-Instruct-Turbo
8B
Smaller
Llama-3.3-70B-Instruct-Turbo
70B
Baseline
Meta-Llama-3.1-405B-Instruct-Turbo
405B
Larger
Qwen
Qwen2.5-7B-Instruct-Turbo
7B
Small
Qwen2.5-72B-Instruct-Turbo
72B
Medium
Qwen3-235B-A22B-Instruct
235B (MoE)
Large
Appendix
Table 7: Complete model configuration across three families. All models accessed via the Together AI API with temperature T=0.2 .
Iter 5–10
Iter 10–20
k
JSm
JSt
JSm
JSt
Jac@ k
TM@ k
R∗
3
0.07±0.04
0.09±0.03
0.06±0.03
0.10±0.03
0.77
54.4–57.6
1–7
5
0.06±0.04
0.12±0.04
0.06±0.03
0.13±0.04
0.77
36.7–40.2
1–14
7
0.06±0.03
0.18±0.05
0.06±0.03
0.18±0.05
0.75
23.9–27.2
1–20
10
0.06±0.03
0.27±0.09
0.07±0.03
0.24±0.06
0.76
11.9–14.9
3–17
15
0.07±0.03
0.23±0.12
0.08±0.03
0.29±0.09
0.72
3.9–5.8
10–18
Appendix
Table 8: Sensitivity of the mode–tail analysis to the cutoff k (US airport travel demand; four primary departure contexts). JSm and JSt are the mode- and tail-region divergences by convergence phase, computed exactly as in Table 2 ; the k=5 row reproduces that table. Divergences are mean ± half-CI over the consecutive-sample comparisons pooled across the four contexts (§ 4 ). Jac@ k : mean top- k Jaccard between consecutive post-convergence samples. TM@ k : percentage of probability mass outside the top- k modes. R∗ : cumulative top- k stabilization point at τ=0.9 . TM@ k and R∗ are given as ranges over the four contexts. Tail dominance holds for k≥5 , where the paired per-comparison difference JSt−JSm excludes zero in both post-convergence phases; at k=3 it holds for iterations 10–20 ( +0.039 , CI [+0.008,0.068] ) but not for 5–10 ( +0.024 , CI [−0.014,0.058] ), which is expected because k=3 truncates the mode region below its natural elbow. Mode agreement is flat in k , whereas R∗ drifts later with k for the mechanical reason given in the text.
Figure 2: Overview of the experimental pipeline. (a) A structured prompt provides the LLM with a traveler persona (age, education), a departure airport, and 300 candidate destination airports, requesting probability assignments as a JSON array. (b) The target LLM generates a response autoregressively at temperature T=0.2 ; this process is repeated R=20 times as independent API calls. (c) Each raw JSON output is parsed and probabilities are renormalized to sum to 1.0. (d) The resulting probability distribution over destinations, shown here for Indianapolis (IND) using real experimental data. High-probability destinations (blue, modes ) and low-probability destinations (red, tail ) are analyzed for distributional consistency, mode–tail dynamics, scale discordance, and prompt sensitivity. See Appendix C.2 for complete prompt templates.
Characteristic
Value
Travel Domain
Choice Set Size
300 airports
Geographic Scope
National (US)
Decision Frequency
Low (occasional)
Price Range
High ($100 and above)
Category Type
Transportation
Appendix
Table 9: Comparison of domain characteristics across experimental settings.
Metric
Travel
Yelp
Amazon
Options
300
300
300
JS converge (iter)
8
11
17
Final support
29
103
51
Top-10 mass (%)
74.7%
49.4%
79.0%
Appendix
Table 10: Cross-domain comparison of convergence metrics across travel demand, Yelp restaurants, and Amazon products. All three are single-context, 20 -iteration convergence runs: travel uses the Indianapolis (IND) context (Table 1 ), matching the single-context Yelp and Amazon experiments; JS converge is the first iteration at which the JS distance to the converged distribution falls below 0.05 , applied identically across domains. Top-10 mass is the probability mass carried by the ten highest-probability items of the converged, renormalized distribution; note the cutoff here is k=10 rather than the k=5 mode boundary used elsewhere (Section 3.4 ), so that all three domains are reported at a single common cutoff. The broader travel result across all 300 departure contexts ( 185 unique destinations) is reported separately in Section F . Referenced from § 5.5 in the main text.
Metric
Original
Rectified
Persona-based Analysis
Success Rate
15% (3/20)
100% (20/20)
Unique Destinations
27
131
Avg Destinations/Persona
13.0
23.9
Total Records
39
477
Convergence Analysis
Appendix
Table 11: Comparison of smaller model (8B) results before and after parsing rectification.
Figure 3: Probability distribution heatmap showing 300 departure airports (rows) × 185 unique destination airports (columns) that appeared across all contexts. Of the 300 airports provided as options, only 185 received a non-zero probability in at least one context. Bright vertical bands reveal consistently high-probability modes.
Figure 4: Cross-domain convergence: Yelp (300 restaurants, top), Amazon (300 products, bottom). Resampling protocol: Section 4 . Each panel plots the mean and 95% band of JS-to-final, Jaccard@5 and support size against the number of samples drawn from the 20 -sample pool. The means are flat across the range while the bands contract, showing that both domains are already stable at small sample counts; the per-iteration convergence points are reported in Table 10 .
Figure 5: Per-sample distributional signatures for the two cross-domain experiments (Yelp restaurants, Amazon products; Indianapolis, 20 iterations; panels and resampling protocol as in Figure 1 ). Cumulative support (panel a) counts canonical items whose averaged, renormalized cumulative probability exceeds 0.1% , the same definition underlying the final-support values in Table 10 ( 103 for Yelp, 51 for Amazon at the full 20 -sample pool). Amazon concentrates probability on a few products (higher maximum probability, smaller support), whereas Yelp spreads mass across many restaurants, consistent with Amazon’s higher top-10 mass ( 79% ) than Yelp’s ( 49% ).
Figure 6: Convergence heatmap for Regional Context A (IND) over 20 iterations. Rows: iterations (0–19). Columns: the distinct destinations observed across the run, i.e. this context’s cumulative support (32 destinations). Bright vertical bands show stable high-probability modes; lighter peripheral colors show gradual tail expansion.
Figure 7: Side-by-side comparison of convergence patterns across three LLM sizes with rectified parsing. Left: Baseline (70B, 32 final destinations); Center: Bigger (405B, 54 final destinations); Right: Smaller (8B, 100 final destinations). The smaller model achieves the highest diversity but requires more iterations for stabilization.
Figure 8: Scale discordance matrix showing pairwise correlations between three LLM sizes (8B, 70B, 405B). Values near zero indicate minimal agreement, demonstrating that model scale does not induce smooth distributional refinement. Different models encode fundamentally different world priors.
Figure 9: Side-by-side comparison of persona-based distributions across three LLM sizes with rectified parsing. Left: Baseline (70B, 46 destinations); Center: Bigger (405B, 72 destinations); Right: Smaller (8B, 131 destinations). The smaller model produces the highest diversity when properly parsed.
Figure 10: Per-sample distributional signatures across three Qwen scales (Indianapolis, 20 samples; Section 4 resampling protocol; panels as in Figure 1 ). The three scales occupy distinct operating points: the largest model (235B, a Mixture-of-Experts) is the most diffuse (highest unique-destination count), while the 7B and 72B models concentrate more probability mass on the top destinations.
Figure 11: Per-sample distributional signatures across three Mistral scales (Indianapolis, 20 samples; Section 4 resampling protocol; panels as in Figure 1 ). The smallest model (Mistral-7B) is by far the most diffuse and least peaked, whereas Mistral-Small-24B and Mistral-8x7B concentrate mass on a few destinations, giving three distinct operating points rather than a refinement of one another.
Figure 12: Persona-based probability distributions from Regional Context A (IND). Rows: 20 demographic personas (Age × Education). Columns: destination outcomes. Personas aged 35 to 44 show the most concentrated preferences and senior personas the most diverse, by both peak probability and entropy.
Airport
JS Dist.
95% CI
FWA (Fort Wayne)
0.359
[0.316, 0.412]
IND (Indianapolis)
0.334
[0.298, 0.375]
ATL (Atlanta)
0.285
[0.255, 0.317]
ORD (Chicago)
0.264
[0.230, 0.306]
Mean
0.311
—
Appendix
Table 12: Seasonality perturbation effect sizes for travel demand (US airports). JS distance (the square root of the JS divergence; Appendix B ) between summer (June–August) and winter (December–February) distributions. Smaller airports show larger seasonal effects than major hubs.
T
Unique
Max P
Entropy
Top-5
Jac@5
Jac@10
0.1
30.5
0.176
3.93
0.573
0.924
0.885
0.2
33.2
0.164
4.07
0.540
0.929
0.911
0.3
37.5
0.160
4.16
0.530
0.924
0.853
0.4
33.3
0.172
4.07
0.550
0.883
0.886
0.5
36.7
0.169
4.14
0.541
0.901
0.890
0.6
35.7
0.175
4.08
0.551
0.873
0.886
Appendix
Table 13: Temperature effects on convergence metrics (mean across 6 settings; Unique, Max P, Entropy and Top-5 at the final iteration, Jac@5 and Jac@10 averaged over all iterations). Bold indicates the primary experimental setting ( T=0.2 ). All ten sampled temperatures are listed. Jac@10 is reported alongside Jac@5 to show that mode-set stability is preserved at both cutoffs across the full temperature range, so the mode–tail threshold of Appendix B.6 is not specific to T=0.2 . Jac@5 declines overall but not monotonically, rising at T=0.1→0.2 , at T=0.4→0.5 , and at T=0.8→0.9 ; the first of these is what makes T=0.2 the highest Jac@5 in the sweep.
Figure 13: Temperature effects across 6 experimental settings (3 airports × 2 personas) at the final iteration. Each panel shows a different metric plotted against temperature ( T∈{0.1,…,1.0} ). Lines represent different setting combinations. Key observations: (1) Unique destinations increase with temperature across most settings; (2) Shannon entropy consistently increases with temperature (6/6 settings); (3) Top-5 mass consistently decreases with temperature (6/6 settings); (4) Jaccard@5 stability peaks at low temperatures and degrades at high temperatures. These patterns confirm that temperature effects are systematic and generalizable across diverse experimental configurations.
Metric (Claim)
T=0.2
T=0
Convergence
IND unique dest, iter 20
32
27
IND stabilization
Iter 3
Iter 2
Scale Discordance
8B unique dest, iter 20
100
100
405B stabilization
Iter 4
Iter 2
Appendix
Table 14: Comparison of key metrics between stochastic ( T=0.2 ) and greedy ( T=0 ) decoding across the three core claims. All qualitative findings are preserved.
Metric
CV Perm.
CV Para.
CV Origins
Unique Dest.
13.2%
14.6%
20.5%
Max Prob.
10.1%
13.6%
13.3%
Entropy
5.1%
4.6%
5.4%
Top-5 Mass
10.1%
9.1%
10.0%
Appendix
Table 15: Coefficient of variation across three robustness sources. All values below 21% indicate good reproducibility.
Metric
Airport
Persona
Model
Max Prob.
−0.02
0.16
0.20
Entropy
−0.02
−0.40
0.17
Top-5 Mass
0.08
−0.07
−0.01
Support
0.07
−0.31
0.46
Appendix
Table 16: ICC(1,1) variance decomposition by grouping factor. A negative value indicates that within-group variance exceeds between-group variance, so the factor accounts for no systematic share. Departure airport stays at or below 0.08 on every metric and persona is at or below zero except on peak probability, whereas model scale is the only factor with a non-trivial positive component and is largest on distributional support.
LLMs are increasingly used to simulate human survey responses, but prior work has mainly evaluated replication using mean-level or aggregate agreement, offering limited insight into whether LLMs reproduce the variability of human behavior. We evaluate LLM-based survey replication at the distributional level using a non-public 2010 consumer choice experiment on Korean instant noodle purchases, a setting unlikely to overlap with model training data. We evaluate three response variables of differing statistical type: binary purchase incidence, categorical brand choice, and count purchase quantity. For each, we compare human and LLM responses at mean-level, pattern, and distributional alignment, and against reference baselines from the human data alone. LLMs reproduce condition-level patterns reasonably well but fail to capture distributional structure: for purchase quantity, no model beats a condition-insensitive baseline that simply matches the pooled human distribution. Because models that match human means well can still produce distributions further from humans than this baseline, mean-based evaluation alone can be actively misleading. Replication also varies with input configuration, with structured personas and multimodal inputs improving alignment while explicit reasoning prompting degrades it monotonically.
We introduce UnpredictaBench, an evaluation that tests the ability of large language models (LLMs) to capture true underlying distributions. As LLMs are increasingly used as substitutes for other entities (e.g., for humans in economic simulations), the tendency of many models to collapse towards a single plausible answer means a failure to capture the unpredictability of real systems. Recent work on improving output diversity is insufficient for this setting: simulation requires samples that are calibrated to a target distribution, not merely varied outputs. UnpredictaBench isolates a simplified but fundamental version of this problem: sampling outcomes from individual target distributions, including canonical statistical distributions, distributions induced by stochastic programs, and natural-language scenarios that describe random processes. We introduce 448 such problems together with KS@N, a general-purpose evaluation metric that quantifies how well a model outputs approximate black-box target distributions via the Kolmogorov-Smirnov statistical test. This is the rate at which we fail to reject model samples of size N against ground-truth samples, with larger N indicating greater difficulty. Tested across open and proprietary models, we find a large spread in distributional capabilities. For instance, when models generate samples of size 100 (KS@100, our standard metric), scores range from near 0 to over 20%. No model is able to achieve over 40% at KS@100, showing significant headroom in distributional sampling as a capability. Although adding reasoning can somewhat increase scores, we find no immediate solution for this issue. UnpredictaBench shows that even simple distributional simulation remains challenging, making it a necessary first step toward using LLMs as stand-ins for complex systems. Project website and resources are available at https://unpredictabenchmark.github.io/.
Amirhossein Abaskohi, Amirhossein Dabiriaghdam, Liang Luo +4
University of British Columbia · 2Independent Researcher
Large language models (LLMs) can readily reproduce conventional expressions, yet their ability to model gradient frequency distributions remains underexplored. We investigate this using linguistic binomials, such as men and women, where both word permutations are grammatically valid but exhibit distinct, cross-linguistic variations in conventionality. We formalize binomial ordering as a distributional alignment problem, and construct a multilingual dataset of 600 binomial pairs across 8 languages. With categorical and distributional metrics, we measure and compare the corpus-derived preferences with model-induced ordering probabilities of 6 open-weight LLMs. While models often behaviorally recover the dominant corpus-preferred order, particularly for strongly conventionalized pairs, they align less well with the exact corpus preference distributions. This suggests that apparent directional order overstates how faithfully LLMs capture the statistical nuances of language use. Sparse probing verifies that the concept of preference strength is partially encoded among middle-to-late layers, and steering along probe-derived directions alters model-induced ordering distributions, demonstrating that the statistical behavioral preference of LLMs can be mechanistically measured and manipulated via internal representations.
Zhiqing Yang, Yilun Liu, Yunpu Ma +2
Ludwig Maximilian University of Munich · Munich Center for Machine Learning