LLM Bidders Preserve the Mechanism-Level Orderings of Human Bidders
Organizations: MIT & NBER
Abstract
Training on vast amounts of human-generated data has motivated growing interest in using large language models (LLMs) to simulate human behavior. We ask which features of human behavior general-purpose models preserve when used out of the box in auctions, where multiple bidders interact under explicit rules and incentives. We evaluate five LLMs across seven laboratory settings against human benchmarks reconstructed from published experiments, with uncertainty bands for the private-value comparisons. Our main focus is on three large models without extended test-time reasoning: GPT-4o, Claude3.5 Haiku, and Gemini2.0 Flash. LLM and human deviations from theory differ in magnitude and often in direction: humans overbid in second-price auctions, whereas most models that deviate underbid. Surprisingly, without task-specific fine-tuning or calibration to human bids, the three non-reasoning large models robustly preserve key orderings of auction formats by deviation from theory. First-price auctions are harder than second-price, and ascending clocks reduce deviations relative to sealed bids wherever data are adequate. Kendall's between the human and GPT-4o difficulty rankings is and positive in every joint bootstrap draw. The reasoning model bids almost at equilibrium in the observed private-value settings, leaving little variation in errors to compare; the small model's large errors yield an inverted ranking. All five models nevertheless reproduce the stronger first-price winner's curse. Clock framing improves bidding for two of the three non-reasoning large models, and GPT-4o recovers the ordering of last-minute bidding across closing rules in an eBay-style marketplace.
Figures & tables
| Format | Human | GPT-4o | Claude 3.5 Haiku | Gemini 2.0 Flash | GPT-5-mini | Llama-3-8B |
|---|---|---|---|---|---|---|
| First-Price CV | 47.59 ∗ | 59.99 (96) | 54.07 (29) † | 59.50 (30) | 30.22 (30) | 53.68 (30) |
| First-Price IPV | 24.76 [22.30, 27.25] | 27.55 (285) | 24.09 (90) | 27.05 (90) | 0.08 (87) | 38.16 (90) |
| Second-Price CV | 18.23 ∗ | 40.60 (92) | 27.19 (29) † | 30.89 (30) | 14.80 (30) | 32.27 (30) |
| Second-Price APV | 9.31 [6.78, 11.74] | 10.73 (291) | 9.58 (90) | 8.02 (87) | 1.36 (87) | 53.08 (90) |
| AC-Closed (AC-B) APV | 5.83 [2.30, 10.11] | 0.47 (114) | 0.00 (2) † | 1.26 (46) | DS ‡ | 67.00 (2) † |
| Second-Price IPV | 5.65 [4.46, 6.79] | 8.69 (270) | 9.72 (90) | 3.01 (87) | 0.20 (87) | 54.89 (90) |
| Model | Formats used ( ) | (all usable) | Joint 95% CI | (3 sealed) | FPSB SPSB | |
|---|---|---|---|---|---|---|
| GPT-4o | all five (5) | 100% | yes (100%) | |||
| Claude 3.5 Haiku | 3 sealed + AC (4) | 100% | yes (100%) | |||
| Gemini 2.0 Flash | 3 sealed + AC-B (4) | 100% | yes (100%) | |||
| GPT-5-mini | 3 sealed (3) | 36% | no (36%) | |||
| Llama-3-8B | 3 sealed (3) | 0% | no (0%) |
| Human regularity | Human | GPT-4o | Claude 3.5 Haiku | Gemini 2.0 Flash | GPT-5-mini | Llama-3-8B | Non-reasoning large models |
| Tier 1: magnitude within the human reconstruction band | |||||||
| FPSB IPV | [22.3, 27.3] | ✓ | ✓ | 2/3 | |||
| SPSB IPV | [4.5, 6.8] | 0/3 | |||||
| Sealed SP-APV | [6.8, 11.7] | ✓ | ✓ | ✓ | 3/3 | ||
| Open clock, AC (APV) | [2.0, 5.4] | n/a | DS ‡ | n/a | 0/2 | ||
| Closed clock, AC-B (APV) | [2.3, 10.1] | n/a | DS ‡ | n/a | 0/2 | ||
| Baseline | Clock framing | within | matches | ||||
|---|---|---|---|---|---|---|---|
| Model | SMAD | mean dev | SMAD | mean dev | baseline clock | (MW) | human sign? |
| GPT-4o | 11.8% | 6.2% | 28% 45% | ||||
| Claude 3.5 Haiku | 8.1% | 7.1% | 28% 7% | ||||
| Gemini 2.0 Flash | 5.0% | 9.3% | 62% 37% | ||||
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Source | Treatment | Matched moment | |||
|---|---|---|---|---|---|
| Li (2017) | 2P | 0.50 | 0.40 | 0.219 | Mean ratio = 1.08 |
| Li (2017) | AC | 0.67 | 0.18 | 0.198 | Mean ratio = 1.02 |
| Breitmoser (2022) | 2P | 0.44 | 0.40 | 0.283 | Mean ratio = 1.10 |
| Breitmoser (2022) | AC | 0.83 | 0.12 | 0.154 | Mean ratio = 1.01 |
| Kagel-Levin (1993) | FPSB, =5 | 0.75 | 0.12 | 0.100 | =0.88 |
| Kagel-Levin (1993) | SPSB, =5 | 0.27 | 0.67 | 0.153 | =0.97 |
| Format | Source | Point | 68% band | 95% band |
|---|---|---|---|---|
| First-Price IPV | Kagel and Levin (1993) | 24.76 | [23.55, 25.97] | [22.30, 27.25] |
| Second-Price IPV | Kagel and Levin (1993) | 5.65 | [5.03, 6.23] | [4.46, 6.79] |
| Third-Price IPV ( ) | Kagel and Levin (1993) | 7.66 | [6.92, 8.39] | [6.15, 9.15] |
| Second-Price APV | Li (2017) | 9.31 | [7.99, 10.56] | [6.78, 11.74] |
| Ascending Clock APV | Li (2017) | 3.54 | [2.69, 4.37] | [2.00, 5.36] |
| AC-Closed (AC-B) APV | Breitmoser and Schweighofer-Kodritsch (2022) | 5.83 | [3.77, 7.75] | [2.30, 10.11] |
| FP-CV | 39.00 (30) | 59.99 (96) | 42.92 (30) |
|---|---|---|---|
| FPSB | 28.07 (300) | 27.55 (285) | 24.99 (297) |
| SP-CV | 35.51 (100) | 40.60 (92) | 33.53 (99) |
| SP-APV | 12.87 (297) | 10.73 (291) | 12.75 (300) |
| AC-B | 0.50 † (4) | 0.47 (114) | 2.15 † (8) |
| SPSB | 14.77 (294) | 8.69 (270) | 12.89 (297) |
| Format | Human | GPT-4o | Claude 3.5 Haiku | Gemini 2.0 Flash | GPT-5-mini | Llama-3-8B |
|---|---|---|---|---|---|---|
| FPSB IPV | 32 / 68 | 14 / 86 | 2 / 98 | 10 / 90 | n/a | 82 / 18 |
| within 6% | within 3%, n=285 | within 1%, n=90 | within 1%, n=90 | within 100%, n=87 (no deviating bids) | within 1%, n=90 | |
| SPSB IPV | 8 / 92 | 81 / 19 | 33 / 67 | 100 / 0 | 100 / 0 | 94 / 6 |
| within 27% | within 31%, n=270 | within 27%, n=90 | within 74%, n=87 | within 99%, n=87 | within 4%, n=90 | |
| SP-APV | 20 / 80 | 92 / 8 | 70 / 30 | 100 / 0 | 100 / 0 | 93 / 7 |
| within 50% | within 9%, n=291 | within 3%, n=90 | within 30%, n=87 | within 83%, n=87 | within 9%, n=90 |
| Auction | Bidders | Mean (Std.) | Median | % Positive | (Profit ) |
|---|---|---|---|---|---|
| FPSB | 4 | (3.87) | 24.0% | ||
| FPSB | 5 | (3.63) | 6.0% | ||
| FPSB | 6 | (3.82) | 4.0% | ||
| FPSB | 7 | (2.93) | 0.0% | ||
| SPSB | 4 | (4.87) | 58.0% | 0.967 | |
| SPSB | 5 | (4.35) | 46.0% | 0.542 |
| Baseline | Clock framing | within | matches | ||||
| Model | SMAD | mean dev | SMAD | mean dev | baseline clock | (MW) | human sign? |
| Gemma-3-27B | 26.3% | 13.5% | 6% 22% | ||||