Authors: Fan Huang, Minsuk Kim, C. Tyler Diggans, Filippo Radicchi
Organizations: Indiana University Bloomington Bloomington, IN, USA · Massachusetts Institute of Technology Cambridge, MA, USA · Root Dynamix LLC Sewickley, PA, USA
Large Language Models (LLMs) are increasingly used as probabilistic generators for simulation, synthetic data generation, and decision support in settings where real-world data are unavailable. Yet, the structure and reliability of the distributions they produce remain understudied. Here, we systematically analyze LLM-generated distributions of preferences for air travel, restaurants, and consumer products. Encouragingly, all models considered in our analysis exhibit self-coherence, with the most probable outcomes stabilizing rapidly under repeated sampling. At the same time, we observe substantial discordance across both model families and scales, with little consensus even among their most probable outcomes. These patterns hold across nine open-weight models, three choice domains, and show robustness under temperature changes, greedy decoding, and perturbations of prompt and ordering. Our findings indicate that outcomes are influenced more by the choice of model than by the wording of the prompt, challenging the common assumption that sufficiently capable LLMs produce similar preference distributions when used as stand-ins for survey respondents.
Figures & tables
Context/Model
Unique@1
Unique@20
Stable@
IND
19
32
Iter 3
FWA
16
42
Iter 6
ATL
15
43
Iter 5
ORD
18
34
Iter 3
Model Size Comparison (IND)
Baseline (70B)
19
32
Iter 3
Table 1: Convergence results for travel demand (US airports). Unique@ R : cumulative support size, distinct destinations with non-zero probability pooled over the first R samples (Section 3.5 ). Stable@ : the first sample from which this cumulative support plateaus (changes by at most two destinations over three successive samples); a support-size stabilization signal, complementary to and distinct from the top- k stabilization point R∗ (Section 3.5 ).
Iter.
JStotal
JSmode
JStail
Δtail−mode
1–3
0.12±0.06
0.07±0.06
0.20±0.07
+0.13±0.03
3–5
0.19±0.09
0.13±0.08
0.31±0.11
+0.18±0.05
5–10
0.08±0.03
0.06±0.04
0.12±0.04
+0.06±0.03
10–20
0.09±0.03
0.06±0.03
0.13±0.04
+0.07±0.04
Table 2: JS divergence by convergence phase on the full distribution, the renormalized top-5 mode region, and its complement (US airport travel demand). Cells are mean ± half-CI (§ 4 ) over consecutive-sample comparisons pooled across the four contexts ( N=12,8,20,34 by phase). Δtail−mode , the paired difference, is the appropriate test of tail dominance since both regions share a sample pair: it excludes zero in every phase, whereas the marginal intervals overlap in three of the four.
Model Pair
JS Div.
Agr.@5
Top1
70B vs 405B
0.36±0.04
0.12±0.05
0.30±0.20
70B vs 8B
0.47±0.05
0.08±0.04
0.05±0.07
405B vs 8B
0.50±0.04
0.08±0.03
0.20±0.18
Table 3: Scale discordance metrics across three Llama sizes (US airport travel demand). Mean ± half-CI over 20 personas; see § 4 for the precision convention.
Figure 1: Per-sample distributional signatures across three LLM sizes (Section 4 resampling protocol; the CI band for the number of sampled distributions equal to 20 is by definition zero because there is only one sample equal to the full pool, so the informative part of each curve is the small- R region). (a) Unique destinations: 8B (mean 71.6) is 1.58 × the 405B (45.2) and 2.45 × the 70B (29.3). (b) Maximum destination probability: 70B and 405B concentrate 0.171 and 0.174, respectively; 8B is more diffuse (0.090). (c) Std of probabilities: 70B and 405B at 0.038; 8B tighter (0.019). The three models occupy structurally different operating points, not refinements of one another: across the 20 personas no two share an identical 70B/405B top-5 (mean top-5 overlap 0.12 ) and even the single dominant destination differs in most cases (top-1 match 0.30 ; Table 3 ), so the larger model re-weights which destinations dominate rather than sharpening the same modes.
Family
S-M
M-L
S-L
Llama
0.47±0.05
0.36±0.04
0.50±0.04
Qwen
0.14 †
0.15 †
0.23 †
Mistral
0.22 †
0.14 †
0.28 †
Table 4: Cross-family scale discordance (JS divergence; US airport travel demand). Llama: mean ± half-CI over 20 personas (same protocol as Table 3 ). † single-context point estimate (IND only); airport names are normalized to IATA codes before comparison (Appendix I ).
Perturbation Type
JS Div.
N
Persona (with vs without)
0.13±0.01
298
Age (within education)
0.33±0.05
24
Education (within age)
0.26±0.03
60
Context (regional vs hub)
0.31±0.04
40
Table 5: Prompt sensitivity effect sizes (JS divergence; US airport travel demand). Mean ± half-CI over N pairwise comparisons; see § 4 for the precision convention.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Method
T
Sup.
Entropy
Top-5
JS
Single-choice
0.2
1
0.00±0.00
1.00±0.00
0.64±0.02
Single-choice
1.0
2
0.18±0.10
1.00±0.00
0.61±0.04
Verbalized
0.2
35
3.97±0.29
0.58±0.06
(ref.)
Verbalized
1.0
34
4.25±0.20
0.46±0.04
(ref.)
Appendix
Table 6: Comparison of distribution elicitation methods for IND departures (Llama-3.3-70B). Sup. : unique destinations observed. Entropy : Shannon entropy in bits. Top-5 : probability mass in top 5 destinations. JS : Jensen–Shannon divergence against the verbalized baseline at the same temperature; “(ref.)” marks the verbalized rows, which are that baseline (self-divergence 0 ). Entropy, Top-5 and JS are mean ± half-CI over 10,000 bootstrap resamples of the underlying calls (20 verbalized calls at T=0.2 and 18 at T=1.0 ; 300 single-choice draws at each). Support is an exact count over the full sample and is given without an interval, because resampling calls can only remove distinct destinations and never add them, so its bootstrap is biased downward. The zero intervals on the single-choice rows are structural rather than rounding: at T=0.2 every draw returned the same destination, and with a support of one or two all mass necessarily falls inside the top five. Single-choice sampling produces degenerate distributions at both temperatures, while verbalized elicitation captures rich distributional structure in a single query.
Family
Model ID
Params
Role
Llama
Llama-3.1-8B-Instruct-Turbo
8B
Smaller
Llama-3.3-70B-Instruct-Turbo
70B
Baseline
Meta-Llama-3.1-405B-Instruct-Turbo
405B
Larger
Qwen
Qwen2.5-7B-Instruct-Turbo
7B
Small
Qwen2.5-72B-Instruct-Turbo
72B
Medium
Qwen3-235B-A22B-Instruct
235B (MoE)
Large
Appendix
Table 7: Complete model configuration across three families. All models accessed via the Together AI API with temperature T=0.2 .
Iter 5–10
Iter 10–20
k
JSm
JSt
JSm
JSt
Jac@ k
TM@ k
R∗
3
0.07±0.04
0.09±0.03
0.06±0.03
0.10±0.03
0.77
54.4–57.6
1–7
5
0.06±0.04
0.12±0.04
0.06±0.03
0.13±0.04
0.77
36.7–40.2
1–14
7
0.06±0.03
0.18±0.05
0.06±0.03
0.18±0.05
0.75
23.9–27.2
1–20
10
0.06±0.03
0.27±0.09
0.07±0.03
0.24±0.06
0.76
11.9–14.9
3–17
15
0.07±0.03
0.23±0.12
0.08±0.03
0.29±0.09
0.72
3.9–5.8
10–18
Appendix
Table 8: Sensitivity of the mode–tail analysis to the cutoff k (US airport travel demand; four primary departure contexts). JSm and JSt are the mode- and tail-region divergences by convergence phase, computed exactly as in Table 2 ; the k=5 row reproduces that table. Divergences are mean ± half-CI over the consecutive-sample comparisons pooled across the four contexts (§ 4 ). Jac@ k : mean top- k Jaccard between consecutive post-convergence samples. TM@ k : percentage of probability mass outside the top- k modes. R∗ : cumulative top- k stabilization point at τ=0.9 . TM@ k and R∗ are given as ranges over the four contexts. Tail dominance holds for k≥5 , where the paired per-comparison difference JSt−JSm excludes zero in both post-convergence phases; at k=3 it holds for iterations 10–20 ( +0.039 , CI [+0.008,0.068] ) but not for 5–10 ( +0.024 , CI [−0.014,0.058] ), which is expected because k=3 truncates the mode region below its natural elbow. Mode agreement is flat in k , whereas R∗ drifts later with k for the mechanical reason given in the text.
Figure 2: Overview of the experimental pipeline. (a) A structured prompt provides the LLM with a traveler persona (age, education), a departure airport, and 300 candidate destination airports, requesting probability assignments as a JSON array. (b) The target LLM generates a response autoregressively at temperature T=0.2 ; this process is repeated R=20 times as independent API calls. (c) Each raw JSON output is parsed and probabilities are renormalized to sum to 1.0. (d) The resulting probability distribution over destinations, shown here for Indianapolis (IND) using real experimental data. High-probability destinations (blue, modes ) and low-probability destinations (red, tail ) are analyzed for distributional consistency, mode–tail dynamics, scale discordance, and prompt sensitivity. See Appendix C.2 for complete prompt templates.
Characteristic
Value
Travel Domain
Choice Set Size
300 airports
Geographic Scope
National (US)
Decision Frequency
Low (occasional)
Price Range
High ($100 and above)
Category Type
Transportation
Appendix
Table 9: Comparison of domain characteristics across experimental settings.
Metric
Travel
Yelp
Amazon
Options
300
300
300
JS converge (iter)
8
11
17
Final support
29
103
51
Top-10 mass (%)
74.7%
49.4%
79.0%
Appendix
Table 10: Cross-domain comparison of convergence metrics across travel demand, Yelp restaurants, and Amazon products. All three are single-context, 20 -iteration convergence runs: travel uses the Indianapolis (IND) context (Table 1 ), matching the single-context Yelp and Amazon experiments; JS converge is the first iteration at which the JS distance to the converged distribution falls below 0.05 , applied identically across domains. Top-10 mass is the probability mass carried by the ten highest-probability items of the converged, renormalized distribution; note the cutoff here is k=10 rather than the k=5 mode boundary used elsewhere (Section 3.4 ), so that all three domains are reported at a single common cutoff. The broader travel result across all 300 departure contexts ( 185 unique destinations) is reported separately in Section F . Referenced from § 5.5 in the main text.
Metric
Original
Rectified
Persona-based Analysis
Success Rate
15% (3/20)
100% (20/20)
Unique Destinations
27
131
Avg Destinations/Persona
13.0
23.9
Total Records
39
477
Convergence Analysis
Appendix
Table 11: Comparison of smaller model (8B) results before and after parsing rectification.
Figure 3: Probability distribution heatmap showing 300 departure airports (rows) × 185 unique destination airports (columns) that appeared across all contexts. Of the 300 airports provided as options, only 185 received a non-zero probability in at least one context. Bright vertical bands reveal consistently high-probability modes.
Figure 4: Cross-domain convergence: Yelp (300 restaurants, top), Amazon (300 products, bottom). Resampling protocol: Section 4 . Each panel plots the mean and 95% band of JS-to-final, Jaccard@5 and support size against the number of samples drawn from the 20 -sample pool. The means are flat across the range while the bands contract, showing that both domains are already stable at small sample counts; the per-iteration convergence points are reported in Table 10 .
Figure 5: Per-sample distributional signatures for the two cross-domain experiments (Yelp restaurants, Amazon products; Indianapolis, 20 iterations; panels and resampling protocol as in Figure 1 ). Cumulative support (panel a) counts canonical items whose averaged, renormalized cumulative probability exceeds 0.1% , the same definition underlying the final-support values in Table 10 ( 103 for Yelp, 51 for Amazon at the full 20 -sample pool). Amazon concentrates probability on a few products (higher maximum probability, smaller support), whereas Yelp spreads mass across many restaurants, consistent with Amazon’s higher top-10 mass ( 79% ) than Yelp’s ( 49% ).
Figure 6: Convergence heatmap for Regional Context A (IND) over 20 iterations. Rows: iterations (0–19). Columns: the distinct destinations observed across the run, i.e. this context’s cumulative support (32 destinations). Bright vertical bands show stable high-probability modes; lighter peripheral colors show gradual tail expansion.
Figure 7: Side-by-side comparison of convergence patterns across three LLM sizes with rectified parsing. Left: Baseline (70B, 32 final destinations); Center: Bigger (405B, 54 final destinations); Right: Smaller (8B, 100 final destinations). The smaller model achieves the highest diversity but requires more iterations for stabilization.
Figure 8: Scale discordance matrix showing pairwise correlations between three LLM sizes (8B, 70B, 405B). Values near zero indicate minimal agreement, demonstrating that model scale does not induce smooth distributional refinement. Different models encode fundamentally different world priors.
Figure 9: Side-by-side comparison of persona-based distributions across three LLM sizes with rectified parsing. Left: Baseline (70B, 46 destinations); Center: Bigger (405B, 72 destinations); Right: Smaller (8B, 131 destinations). The smaller model produces the highest diversity when properly parsed.
Figure 10: Per-sample distributional signatures across three Qwen scales (Indianapolis, 20 samples; Section 4 resampling protocol; panels as in Figure 1 ). The three scales occupy distinct operating points: the largest model (235B, a Mixture-of-Experts) is the most diffuse (highest unique-destination count), while the 7B and 72B models concentrate more probability mass on the top destinations.
Figure 11: Per-sample distributional signatures across three Mistral scales (Indianapolis, 20 samples; Section 4 resampling protocol; panels as in Figure 1 ). The smallest model (Mistral-7B) is by far the most diffuse and least peaked, whereas Mistral-Small-24B and Mistral-8x7B concentrate mass on a few destinations, giving three distinct operating points rather than a refinement of one another.
Figure 12: Persona-based probability distributions from Regional Context A (IND). Rows: 20 demographic personas (Age × Education). Columns: destination outcomes. Personas aged 35 to 44 show the most concentrated preferences and senior personas the most diverse, by both peak probability and entropy.
Airport
JS Dist.
95% CI
FWA (Fort Wayne)
0.359
[0.316, 0.412]
IND (Indianapolis)
0.334
[0.298, 0.375]
ATL (Atlanta)
0.285
[0.255, 0.317]
ORD (Chicago)
0.264
[0.230, 0.306]
Mean
0.311
—
Appendix
Table 12: Seasonality perturbation effect sizes for travel demand (US airports). JS distance (the square root of the JS divergence; Appendix B ) between summer (June–August) and winter (December–February) distributions. Smaller airports show larger seasonal effects than major hubs.
T
Unique
Max P
Entropy
Top-5
Jac@5
Jac@10
0.1
30.5
0.176
3.93
0.573
0.924
0.885
0.2
33.2
0.164
4.07
0.540
0.929
0.911
0.3
37.5
0.160
4.16
0.530
0.924
0.853
0.4
33.3
0.172
4.07
0.550
0.883
0.886
0.5
36.7
0.169
4.14
0.541
0.901
0.890
0.6
35.7
0.175
4.08
0.551
0.873
0.886
Appendix
Table 13: Temperature effects on convergence metrics (mean across 6 settings; Unique, Max P, Entropy and Top-5 at the final iteration, Jac@5 and Jac@10 averaged over all iterations). Bold indicates the primary experimental setting ( T=0.2 ). All ten sampled temperatures are listed. Jac@10 is reported alongside Jac@5 to show that mode-set stability is preserved at both cutoffs across the full temperature range, so the mode–tail threshold of Appendix B.6 is not specific to T=0.2 . Jac@5 declines overall but not monotonically, rising at T=0.1→0.2 , at T=0.4→0.5 , and at T=0.8→0.9 ; the first of these is what makes T=0.2 the highest Jac@5 in the sweep.
Figure 13: Temperature effects across 6 experimental settings (3 airports × 2 personas) at the final iteration. Each panel shows a different metric plotted against temperature ( T∈{0.1,…,1.0} ). Lines represent different setting combinations. Key observations: (1) Unique destinations increase with temperature across most settings; (2) Shannon entropy consistently increases with temperature (6/6 settings); (3) Top-5 mass consistently decreases with temperature (6/6 settings); (4) Jaccard@5 stability peaks at low temperatures and degrades at high temperatures. These patterns confirm that temperature effects are systematic and generalizable across diverse experimental configurations.
Metric (Claim)
T=0.2
T=0
Convergence
IND unique dest, iter 20
32
27
IND stabilization
Iter 3
Iter 2
Scale Discordance
8B unique dest, iter 20
100
100
405B stabilization
Iter 4
Iter 2
Appendix
Table 14: Comparison of key metrics between stochastic ( T=0.2 ) and greedy ( T=0 ) decoding across the three core claims. All qualitative findings are preserved.
Metric
CV Perm.
CV Para.
CV Origins
Unique Dest.
13.2%
14.6%
20.5%
Max Prob.
10.1%
13.6%
13.3%
Entropy
5.1%
4.6%
5.4%
Top-5 Mass
10.1%
9.1%
10.0%
Appendix
Table 15: Coefficient of variation across three robustness sources. All values below 21% indicate good reproducibility.
Metric
Airport
Persona
Model
Max Prob.
−0.02
0.16
0.20
Entropy
−0.02
−0.40
0.17
Top-5 Mass
0.08
−0.07
−0.01
Support
0.07
−0.31
0.46
Appendix
Table 16: ICC(1,1) variance decomposition by grouping factor. A negative value indicates that within-group variance exceeds between-group variance, so the factor accounts for no systematic share. Departure airport stays at or below 0.08 on every metric and persona is at or below zero except on peak probability, whereas model scale is the only factor with a non-trivial positive component and is largest on distributional support.