Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty
Organizations: Hong Kong University of Science and Technology · Huazhong University of Science and Technology · University of Illinois Urbana-Champaign
Abstract
Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty. Although recent studies have developed frameworks to estimate PT parameters for Large Language Models (LLMs), few have examined whether PT itself adequately describes LLM decision-making behavior. To address these gaps, we develop a streamlined workflow grounded in a classic behavioral economics experimental paradigm. First, we estimate PT parameters and evaluate how well the resulting model captures LLM decision-making behavior. We then derive probability mappings for epistemic markers in the same context and inject them into prompts to examine the stability of PT parameters under linguistic uncertainty. Our findings suggest that PT does not consistently provide a reliable account of LLM decision-making across models, and that its application to LLMs is likely sensitive to epistemic uncertainty. The findings caution against the deployment of PT-based frameworks in real-world applications where epistemic ambiguity is prevalent, giving valuable insights in behaviour interpretation and future alignment direction for LLM decision-making.
Figures & tables
| No. | Epistemic Marker | Probability Mapping by Human |
| 1 | almost certain | 95% |
| 2 | highly likely | 90% |
| 3 | very likely | 90% |
| 4 | likely | 80% |
| 5 | probable | 70% |
| 6 | somewhat likely | 70% |
| Model | MAE | R 2 | |||
| Human | 0.670 | 2.630 | 0.685 | - | - |
| Llama-3.1-8B-Instruct | 0.585 | 0.010 | 0.753 | 0.332 | 0.092 |
| Mistral-7B-Instruct-v0.3 | 0.534 | 0.570 | 0.577 | 0.155 | 0.132 |
| Qwen2.5-7B-Instruct | 0.429 | 0.010 | 3.645 | 0.047 | 0.116 |
| Qwen2.5-14B-Instruct | 0.503 | 1.909 | 0.896 | 0.257 | 0.067 |
| Qwen2.5-32B-Instruct | 0.598 | 1.213 | 0.867 | 0.161 | 0.225 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | Value |
| Temperature | 0.7 |
| Maximum new tokens | 8 |
| History length | 10 |
| Number of lottery rounds | 35 |
| Setup | Description | Value |
| Number of repeated samples per lottery question | ||
| Number of lottery-choice questions | ||
| Significance level for confidence intervals | ||
| Confidence level |
| Model | 30% | 70% | 10% | 90% |
| Qwen2.5-7B-Instruct | uncertain | almost certain | somewhat likely | highly likely |
| Llama3.1-8B-Instruct | likely | almost certain | very unlikely | almost certain |
| Mistral-7B-Instruct-v0.3 | very unlikely | highly likely | highly unlikely | almost certain |
| Qwen2.5-14B-Instruct | somewhat unlikely | highly likely | very unlikely | almost certain |
| Qwen2.5-32B-Instruct | somewhat unlikely | probable | somewhat likely | almost certain |
| Option K | Option U | |||
| Lottery | 30% | 70% | 10% | 90% |
| 1 | 40 | 10 | 68 | 5 |
| 2 | 40 | 10 | 75 | 5 |
| 3 | 40 | 10 | 83 | 5 |
| 4 | 40 | 10 | 93 | 5 |
| 5 | 40 | 10 | 106 | 5 |
| Option K | Option U | |||
| Lottery | 90% | 10% | 70% | 30% |
| 1 | 40 | 30 | 54 | 5 |
| 2 | 40 | 30 | 56 | 5 |
| 3 | 40 | 30 | 58 | 5 |
| 4 | 40 | 30 | 60 | 5 |
| 5 | 40 | 30 | 62 | 5 |
| Option K | Option U | |||
| 50% | 50% | 50% | 50% | |
| Lottery | Win | Lose | Win | Lose |
| 1 | 25 | 4 | 30 | 21 |
| 2 | 4 | 4 | 30 | 21 |
| 3 | 1 | 4 | 30 | 21 |
| 4 | 1 | 4 | 30 | 16 |
| Model | Top 7 Epistemic Markers | ||||||
| almost certain | highly likely | very likely | likely | probable | somewhat likely | possible | |
| Llama-3.1-8B-Instruct | 87.92 | 56.04 | 58.00 | 41.80 | 44.23 | 36.71 | 36.29 |
| Mistral-7B-Instruct-v0.3 | 96.80 | 67.89 | 63.10 | 57.50 | 87.22 | 48.98 | 52.16 |
| Qwen2.5-7B-Instruct | 82.71 | 67.00 | 67.06 | 0 4.78 | 0 3.44 | 0 8.93 | 0 4.51 |
| Qwen2.5-14B-Instruct | 91.56 | 55.00 | 54.10 | 42.38 | 26.51 | 32.47 | 38.38 |
| Qwen2.5-32B-Instruct | 97.50 | 95.08 | 82.82 | 65.00 | 54.74 | 46.08 | 55.00 |
| Model | Bottom 7 Epistemic Markers | ||||||
| uncertain | somewhat unlikely | unlikely | not likely | doubtful | very unlikely | highly unlikely | |
| Llama-3.1-8B-Instruct | 35.49 | 33.71 | 33.49 | 36.91 | 33.65 | 31.30 | 32.83 |
| Mistral-7B-Instruct-v0.3 | 48.08 | 40.15 | 38.27 | 30.93 | 34.03 | 29.47 | 27.88 |
| Qwen2.5-7B-Instruct | 35.58 | 27.24 | 19.90 | 27.70 | 25.69 | 18.37 | 19.32 |
| Qwen2.5-14B-Instruct | 29.03 | 26.45 | 19.10 | 20.82 | 13.04 | 10.94 | 10.52 |
| Qwen2.5-32B-Instruct | 0 2.98 | 21.89 | 0 3.33 | 0 3.08 | 0 3.42 | 0 2.77 | 0 2.51 |
| Model | (95% CI) | ||||
| baseline | round1 | round2 | round3 | round4 | |
| Llama-3.1-8B-Instruct | |||||
| Mistral-7B-Instruct-v0.3 | |||||
| Qwen2.5-7B-Instruct | |||||
| Qwen2.5-14B-Instruct | |||||
| Qwen2.5-32B-Instruct | |||||
| Model | (95% CI) | ||||
| baseline | round1 | round2 | round3 | round4 | |
| Llama-3.1-8B-Instruct | |||||
| Mistral-7B-Instruct-v0.3 | |||||
| Qwen2.5-7B-Instruct | |||||
| Qwen2.5-14B-Instruct | |||||
| Qwen2.5-32B-Instruct | |||||
| Model | (95% CI) | ||||
| baseline | round1 | round2 | round3 | round4 | |
| Llama-3.1-8B-Instruct | |||||
| Mistral-7B-Instruct-v0.3 | |||||
| Qwen2.5-7B-Instruct | |||||
| Qwen2.5-14B-Instruct | |||||
| Qwen2.5-32B-Instruct | |||||
| Model | MAE | McFadden | ||||||||
| baseline | round1 | round2 | round3 | round4 | baseline | round1 | round2 | round3 | round4 | |
| Llama-3.1-8B-Instruct | 0.332 | 0.320 | 0.326 | 0.242 | 0.155 | 0.092 | 0.121 | 0.090 | 0.082 | 0.335 |
| Mistral-7B-Instruct-v0.3 | 0.155 | 0.202 | 0.196 | 0.067 | 0.075 | 0.132 | 0.162 | 0.190 | 0.056 | 0.130 |
| Qwen2.5-7B-Instruct | 0.047 | 0.125 | 0.159 | 0.387 | 0.157 | 0.116 | 0.000 | 0.011 | -0.001 | 0.032 |
| Qwen2.5-14B-Instruct | 0.257 | 0.152 | 0.069 | 0.150 | 0.157 | 0.067 | 0.070 | 0.075 | 0.227 | 0.336 |
| Qwen2.5-32B-Instruct | 0.161 | 0.127 | 0.129 | 0.224 | 0.256 | 0.225 | 0.195 | 0.211 | 0.139 | 0.088 |