Training on vast amounts of human-generated data has motivated growing interest in using large language models (LLMs) to simulate human behavior. We ask which features of human behavior general-purpose models preserve when used out of the box in auctions, where multiple bidders interact under explicit rules and incentives. We evaluate five LLMs across seven laboratory settings against human benchmarks reconstructed from published experiments, with uncertainty bands for the private-value comparisons. Our main focus is on three large models without extended test-time reasoning: GPT-4o, Claude3.5 Haiku, and Gemini2.0 Flash. LLM and human deviations from theory differ in magnitude and often in direction: humans overbid in second-price auctions, whereas most models that deviate underbid. Surprisingly, without task-specific fine-tuning or calibration to human bids, the three non-reasoning large models robustly preserve key orderings of auction formats by deviation from theory. First-price auctions are harder than second-price, and ascending clocks reduce deviations relative to sealed bids wherever data are adequate. Kendall's τb between the human and GPT-4o difficulty rankings is 0.60 and positive in every joint bootstrap draw. The reasoning model bids almost at equilibrium in the observed private-value settings, leaving little variation in errors to compare; the small model's large errors yield an inverted ranking. All five models nevertheless reproduce the stronger first-price winner's curse. Clock framing improves bidding for two of the three non-reasoning large models, and GPT-4o recovers the ordering of last-minute bidding across closing rules in an eBay-style marketplace.
Figures & tables
Figure 1 : Human vs. LLM bidding behavior across seven auction formats, non-reasoning large models. Scaled mean absolute deviation (SMAD) from the theoretical benchmark for the reconstructed human data (blue; black whiskers are 95% reconstruction bands from the parametric bootstrap of Appendix A.2 ; common-value anchors have no band) and for the three non-reasoning large models GPT-4o, Claude 3.5 Haiku, and Gemini 2.0 Flash (95% confidence intervals), with N=3 bidders. Hatched bars have fewer than 30 usable bids. Formats are ordered by human SMAD. The reasoning model GPT-5-mini and the small model Llama-3-8B are reported in Table 1 and in Figure 6 .
Format
Human
GPT-4o
Claude 3.5 Haiku
Gemini 2.0 Flash
GPT-5-mini
Llama-3-8B
First-Price CV
47.59 ∗
59.99 (96)
54.07 (29) †
59.50 (30)
30.22 (30)
53.68 (30)
First-Price IPV
24.76 [22.30, 27.25]
27.55 (285)
24.09 (90)
27.05 (90)
0.08 (87)
38.16 (90)
Second-Price CV
18.23 ∗
40.60 (92)
27.19 (29) †
30.89 (30)
14.80 (30)
32.27 (30)
Second-Price APV
9.31 [6.78, 11.74]
10.73 (291)
9.58 (90)
8.02 (87)
1.36 (87)
53.08 (90)
AC-Closed (AC-B) APV
5.83 [2.30, 10.11]
0.47 (114)
0.00 (2) †
1.26 (46)
DS ‡
67.00 (2) †
Second-Price IPV
5.65 [4.46, 6.79]
8.69 (270)
9.72 (90)
3.01 (87)
0.20 (87)
54.89 (90)
Table 1: SMAD (%) by auction format and model. Human anchors are moment-matched reconstructions; brackets give 95% reconstruction bands ( ∗ : common-value anchors are profit-based moments and carry no band). LLM cells report SMAD with the number of usable bids in parentheses; † marks cells with fewer than 30 bids. DS ‡ marks the two ascending-clock cells of GPT-5-mini, in which the model states the dominant strategy outright and does not complete the tick-by-tick clock procedure in any run, so no bids are recorded; these cells enter no statistic. All models run at temperature 0.5 (ignored by GPT-5-mini). We do not report significance tests of LLM cells against the reconstructed anchors, which are reconstructions of published moments rather than samples.
Figure 2 : Direction of deviating bids by model. Shares of underbids (red) and overbids (blue) after removing bids within ±2% of the theoretical benchmark, for human bidders and each model, by auction format. Human shares are the frequencies reported by the source studies relative to value for the truthful formats, and the moment-matched reconstruction relative to the equilibrium bid for first-price. Hatched bars have fewer than 30 usable bids; in the clock formats GPT-5-mini states the dominant strategy and records no bids.
Model
Formats used ( n≥30 )
τb (all usable)
Joint 95% CI
P(τb>0)
τb (3 sealed)
FPSB > SPSB
GPT-4o
all five (5)
+0.60
[+0.40,+1.00]
100%
+1.00
yes (100%)
Claude 3.5 Haiku
3 sealed + AC (4)
+0.67
[+0.67,+1.00]
100%
+0.33
yes (100%)
Gemini 2.0 Flash
3 sealed + AC-B (4)
+0.67
[+0.33,+1.00]
100%
+1.00
yes (100%)
GPT-5-mini
3 sealed (3)
−0.33
[−0.33,+0.33]
36%
−0.33
no (36%)
Llama-3-8B
3 sealed (3)
−1.00
[−1.00,−0.33]
0%
−1.00
no (0%)
Table 2: Kendall’s τb between the human and each model’s SMAD ranking over the private-value formats, with joint 95% bootstrap intervals over LLM bid resampling and human reconstruction re-draws ( B=2000 ). GPT-5-mini’s clock cells record no bids (Table 1 ) and are excluded.
Human regularity
Human
GPT-4o
Claude 3.5 Haiku
Gemini 2.0 Flash
GPT-5-mini
Llama-3-8B
Non-reasoning large models
Tier 1: magnitude within the human reconstruction band
FPSB IPV
[22.3, 27.3]
×
✓
✓
×
×
2/3
SPSB IPV
[4.5, 6.8]
×
×
×
×
×
0/3
Sealed SP-APV
[6.8, 11.7]
✓
✓
✓
×
×
3/3
Open clock, AC (APV)
[2.0, 5.4]
×
×
n/a
DS ‡
n/a
0/2
Closed clock, AC-B (APV)
[2.3, 10.1]
×
n/a
×
DS ‡
n/a
0/2
Table 3: Scorecard by tier. Rows are features of human behavior; cells record whether a model reproduces them. Tier 1: SMAD inside (✓) or outside ( × ) the human 95% reconstruction band, which the human column reports. Tiers 2 and 3: ✓reproduced, × not reproduced, ≈ uninformative because the cells sit at the theoretical benchmark; the human column gives the share of reconstruction re-draws in which the ordering holds where applicable. n/a: fewer than 30 usable bids; a : 46 usable bids; DS ‡ : GPT-5-mini states the dominant strategy outright and does not complete the clock procedure, so no clock bids are recorded. The last column counts the non-reasoning large models that reproduce the row, out of those with data.
Baseline
Clock framing
within ±2%
matches
Model
SMAD
mean dev
SMAD
mean dev
baseline → clock
Δ (MW)
human sign?
GPT-4o
11.8%
−0.118
6.2%
−0.062
28% → 45%
↓p<0.001
√
Claude 3.5 Haiku
8.1%
+0.022
7.1%
−0.042
28% → 7%
↓p<0.001
√
Gemini 2.0 Flash
5.0%
−0.050
9.3%
−0.093
62% → 37%
↑p<0.001
×
Table 4: Clock framing of the second-price sealed-bid auction across the three non-reasoning large model families. SMAD (%), mean scaled deviation (b−v)/E[b⋆] , and the share of bids within ±2% of value under the baseline description (two pooled runs of the standard description) and under the clock framing; p -values are Mann–Whitney tests against the same model’s baseline. The final column records whether the sign of the effect matches the human evidence of Breitmoser and Schweighofer-Kodritsch (2022) : √ if SMAD falls with p<0.05 , × otherwise.
Figure 3 : Final winning-bid timing in simulated eBay-style auctions, IPV environment: standard hard close (left) vs. soft-close extension rule (right). The final winning bid is the last period in which the eventual winner raises her maximum bid; the dashed line marks the scheduled close. The soft-close rule moves the timing of the final winning bid substantially earlier.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Source
Treatment
peq
pover
λ
Matched moment
Li (2017)
2P
0.50
0.40
0.219
Mean ratio = 1.08
Li (2017)
AC
0.67
0.18
0.198
Mean ratio = 1.02
Breitmoser (2022)
2P
0.44
0.40
0.283
Mean ratio = 1.10
Breitmoser (2022)
AC
0.83
0.12
0.154
Mean ratio = 1.01
Kagel-Levin (1993)
FPSB, n =5
0.75
0.12
0.100
R2 =0.88
Kagel-Levin (1993)
SPSB, n =5
0.27
0.67
0.153
R2 =0.97
Appendix
Table 5: Fitted mixture model parameters by source and treatment. Mixture weights from reported rates; λ chosen to match the secondary statistic listed; punder=1−peq−pover .
Format
Source
Point
68% band
95% band
First-Price IPV
Kagel and Levin (1993)
24.76
[23.55, 25.97]
[22.30, 27.25]
Second-Price IPV
Kagel and Levin (1993)
5.65
[5.03, 6.23]
[4.46, 6.79]
Third-Price IPV ( n=5 )
Kagel and Levin (1993)
7.66
[6.92, 8.39]
[6.15, 9.15]
Second-Price APV
Li (2017)
9.31
[7.99, 10.56]
[6.78, 11.74]
Ascending Clock APV
Li (2017)
3.54
[2.69, 4.37]
[2.00, 5.36]
AC-Closed (AC-B) APV
Breitmoser and Schweighofer-Kodritsch (2022)
5.83
[3.77, 7.75]
[2.30, 10.11]
Appendix
Table 6 : Reconstruction-uncertainty bands on the human SMAD anchors (parametric bootstrap over the fitted mixture parameters, B=2000 ; 68% and 95% percentile bands). Common-value anchors have no reconstruction band because they are profit-based moments taken directly from the source.
Figure 4 : Multiple-round simulation of First-Price and Second-Price sealed-bid auction with GPT-4o model (Temp=0.5). Y-axis stands for Scaled Mean Absolute Deviation (SMAD) between actual bidding and optimal bidding.
Figure 5 : Comparison of the bidding behavior of human participants and GPT-4o across three different temperature settings (T=0.1, 0.5, 1.0) in seven auction formats. The horizontal bars represent the Scaled Mean Absolute Deviation (SMAD) from theoretical equilibrium, with error bars indicating 95% confidence intervals.
T=0.1
T=0.5
T=1.0
FP-CV
39.00 (30)
59.99 (96)
42.92 (30)
FPSB
28.07 (300)
27.55 (285)
24.99 (297)
SP-CV
35.51 (100)
40.60 (92)
33.53 (99)
SP-APV
12.87 (297)
10.73 (291)
12.75 (300)
AC-B
0.50 † (4)
0.47 (114)
2.15 † (8)
SPSB
14.77 (294)
8.69 (270)
12.89 (297)
Appendix
Table 7 : Temperature effects on GPT-4o bidding. Top: SMAD (%) by format with the number of usable bids (common value: auction rounds) in parentheses; † : fewer than 30 usable bids; n/a: no data. Bottom: the ordering and direction statistics at each temperature; Kendall’s τb is against the reconstructed human ranking with joint 95% bootstrap intervals.
Figure 6 : Bidding behavior of human participants (reconstructed; whiskers are 95% reconstruction bands) and of the five LLMs (GPT-4o, Claude 3.5 Haiku, Gemini 2.0 Flash, GPT-5-mini, Llama-3-8B) in seven auction formats. Bars are SMAD from the theoretical benchmark with 95% confidence intervals for LLM cells; cells with fewer than 30 usable bids are marked; in the clock formats GPT-5-mini states the dominant strategy and records no bids.
Format
Human
GPT-4o
Claude 3.5 Haiku
Gemini 2.0 Flash
GPT-5-mini
Llama-3-8B
FPSB IPV
32 / 68
14 / 86
2 / 98
10 / 90
n/a
82 / 18
within 6%
within 3%, n=285
within 1%, n=90
within 1%, n=90
within 100%, n=87 (no deviating bids)
within 1%, n=90
SPSB IPV
8 / 92
81 / 19
33 / 67
100 / 0
100 / 0
94 / 6
within 27%
within 31%, n=270
within 27%, n=90
within 74%, n=87
within 99%, n=87
within 4%, n=90
SP-APV
20 / 80
92 / 8
70 / 30
100 / 0
100 / 0
93 / 7
within 50%
within 9%, n=291
within 3%, n=90
within 30%, n=87
within 83%, n=87
within 9%, n=90
Appendix
Table 8: Direction of deviating bids by model and format: percentage of underbids / overbids among bids outside the ±2% band around the theoretical benchmark, with the within-band share and n below. † : fewer than 30 usable bids; DS ‡ : GPT-5-mini states the dominant strategy outright and does not complete the clock procedure; no bids recorded.
Auction
Bidders
Mean (Std.)
Median
% Positive
p (Profit <0 )
FPSB
4
−2.54 (3.87)
−3.00
24.0%
<0.001∗∗∗
FPSB
5
−4.40 (3.63)
−4.00
6.0%
<0.001∗∗∗
FPSB
6
−5.46 (3.82)
−5.05
4.0%
<0.001∗∗∗
FPSB
7
−5.55 (2.93)
−5.30
0.0%
<0.001∗∗∗
SPSB
4
1.30 (4.87)
1.00
58.0%
0.967
SPSB
5
0.07 (4.35)
0.00
46.0%
0.542
Appendix
Table 9 : Winner profit in common-value auctions with LLM bidders. The last column reports one-sided p -values for the hypothesis that winner profit is negative. *** denotes p<0.001 .
Baseline
Clock framing
within ±2%
matches
Model
SMAD
mean dev
SMAD
mean dev
baseline → clock
Δ (MW)
human sign?
Gemma-3-27B
26.3%
−0.261
13.5%
−0.112
6% → 22%
↓p<0.001
√
Appendix
Table 10: Clock framing of the second-price sealed-bid auction for Gemma-3-27B. Columns as in Table 4 .
Figure 7 : Clock framing of the second-price sealed-bid auction for Gemma-3-27B: distribution of per-bid scaled deviations (b−v)/E[b⋆] under the baseline description (grey) and the clock framing (blue), with means marked.
Figure 8 : CDF of final winning-bid timing under the two closing rules, eBay-style auctions, IPV environment. The cumulative distribution function of the period of the final winning bid under the standard hard close (T1) and the soft-close extension rule (T2). The soft close shifts the final winning bid substantially earlier, reproducing the ordering of last-minute bidding across closing rules documented for eBay and early Amazon auctions ( Roth and Ockenfels, 2002 ; Ockenfels and Roth, 2006 ) .