HouseholdBench: Evaluating Large Language Models as Predictors of Household Economic Behavior
Organizations: University of Michigan, Ann Arbor · London School of Economics · MobLab Inc · University of California, Berkeley
Abstract
Large language models (LLMs) have the potential to meet a key goal in economics: a quantitative model of household decision making, across a variety of settings. Yet existing evaluations cover few surveys and outcomes, and do not study how households adjust to changing economic conditions. We introduce a new evaluation, HouseholdBench, which unites 6 U.S. household surveys and 32 prediction tasks spanning numeric, categorical and probabilistic outcomes, related to consumption, income, labor, expectations, and housing. Using past behavior, demographics and macroeconomic conditions, the tasks test whether LLMs predict behavior, including how households adjust to changes in various policies. We evaluate 13 proprietary and open-weight LLMs against a no-change baseline and a gradient-boosted tree model. Most LLMs outperform the no-change baseline, including for policy response tasks -- with the best model lowering error for numeric outcomes by 12.2%. Across most tasks, gradient-boosted trees rank first; leading proprietary LLMs approach their performance, but open-weight models lag. LLMs exhibit systematic over- and underprediction across different tasks. We identify methods that enable a 4 billion parameter open-weight model to match proprietary models' performance: fine-tuning and aggregating 16 predictions per observation. Improvements generalize to policy-response tasks, which are excluded from fine-tuning. We release our datasets, code, and leaderboard on our website: https://jn-huang.github.io/householdbench
Figures & tables
| Baseline tasks | Policy-response tasks | ||||
|---|---|---|---|---|---|
| Task ID | Prediction target | Type | Task ID | Prediction target | Type |
| Consumption and saving | |||||
| cons_cex_categories | Spending in 12 categories | N | cons_cex_rebate01 | Spending after the 2001 tax rebates | N |
| cons_cex_total | Total, nondurable, and durable spending | N | cons_cex_sspay | Daily spending around Social Security payments | N |
| cons_psid_wealth | Net worth in the next wave | N | cons_cex_stimulus08 | Spending after the 2008 stimulus payments | N |
| cons_sce_growth | Year-on-year growth in monthly spending | N | cons_psid_jobloss | Food spending the year after a job loss | N |
| Study | Data | Years | Outcomes predicted | Design | ||||
|---|---|---|---|---|---|---|---|---|
| Spending | Income | Labor | Macro | Housing | Policy | |||
| Athey et al. (2026) | US labor panels (PSID, NLSY) | 1979–2021 | ||||||
| Cruz et al. (2024) | American Community Survey (ACS) | 2018 | ||||||
| Brynjolfsson et al. (2025) | PSID, UK valuation surveys | 2017–2021 | ||||||
| Jia et al. (2026) | Dutch household panel (LISS) | 2023–2024 | ||||||
| Wu et al. (2025) | US consumer surveys (SCE, Nielsen) | 2018–2023 | ||||||
| Average rank within topic | Average score by target type | |||||||||
| Rank | Model | Consumption | Housing | Income | Labor | Macroeconomic | Average rank | RelMAE | Macro-F1 | TV |
| (# = 9) | (# = 5) | (# = 5) | (# = 10) | (# = 3) | (# = 18) | (# = 9) | (# = 5) | |||
| 1 | XGBoost | 2.3 (0.6) | 5.8 (1.0) | 2.6 (0.8) | 4.4 (0.7) | 2.0 (0.8) | 3.4 (0.4) | 0.811 (0.011) | 0.432 (0.014) | 0.141 (0.004) |
| [1pt/2pt] 2 | Fable 5.1 | 6.0 (0.5) | 4.4 (0.8) | 2.4 (0.9) | 4.2 (0.6) | 2.0 (0.9) | 3.8 (0.3) | 0.878 (0.013) | 0.413 (0.013) | 0.139 (0.005) |
| 3 | GPT-5.6 sol | 5.3 (0.5) | 5.2 (0.7) | 6.2 (1.0) | 4.8 (0.5) | 6.0 (0.7) | 5.5 (0.3) | 0.910 (0.013) | 0.399 (0.011) | 0.139 (0.005) |
| 4 | GPT-6 astra | 4.4 (0.4) | 8.6 (0.7) | 4.6 (0.7) | 6.4 (0.4) | 5.0 (0.8) | 5.8 (0.3) | 0.915 (0.014) | 0.369 (0.006) | 0.139 (0.005) |
| Baseline tasks (seen in fine-tuning) | Policy-response tasks (held out) | |||||||
| Model | RelMAE | Macro-F1 | TV | RelMAE | Macro-F1 | TV | ||
| (# = 8) | (# = 6) | (# = 4) | (# = 10) | (# = 3) | (# = 1) | |||
| Reference | XGBoost | 1 | 0.84 | 0.46 | 0.15 | 0.79 | 0.36 | 0.12 |
| [1pt/2pt] | GPT-6 astra | 1 | 0.89 | 0.39 | 0.15 | 0.94 | 0.32 | 0.11 |
| Claude Opus 5 | 1 | 0.90 | 0.47 | 0.15 | 0.91 | 0.28 | 0.13 | |
| Fable 5.1 | 1 | 0.89 | 0.44 | 0.15 | 0.86 | 0.35 | 0.11 | |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Baseline tasks (seen in fine-tuning) | Policy-response tasks (held out) | |||||||
| Model | RelMAE | Macro-F1 | TV | RelMAE | Macro-F1 | TV | ||
| (# = 8) | (# = 6) | (# = 4) | (# = 10) | (# = 3) | (# = 1) | |||
| Reference | XGBoost | 1 | 0.84 | 0.46 | 0.15 | 0.79 | 0.36 | 0.12 |
| [1pt/2pt] | GPT-6 astra | 1 | 0.89 | 0.39 | 0.15 | 0.94 | 0.32 | 0.11 |
| Claude Opus 5 | 1 | 0.90 | 0.47 | 0.15 | 0.91 | 0.28 | 0.13 | |
| Fable 5.1 | 1 | 0.89 | 0.44 | 0.15 | 0.86 | 0.35 | 0.11 | |
| Model | Improvement of from (%) | Improvement of from (%) | ||||
|---|---|---|---|---|---|---|
| Numeric , RelMAE (# = 18) | ||||||
| Proprietary | Fable 5.1 | 0.878 | 0.879 | 0.2 | – | – |
| Claude Opus 5 | 0.905 | 0.902 | 0.3 | – | – | |
| GPT-5.6 sol | 0.910 | 0.910 | 0.0 | – | – | |
| [1pt/2pt] Open-weight | Qwen3.5-4B | 1.304 | 1.272 | 2.5 | 1.223 | 6.3 |
| + SFT | 1.164 | 1.087 | 6.6 | 1.017 | 12.6 | |
| Hyperparameter | Value |
|---|---|
| LoRA rank | 8 |
| LoRA alpha | 32 |
| LoRA dropout | 0.05 |
| LoRA target modules | all linear layers of the language model |
| Optimizer | AdamW ( , , weight decay ) |
| Peak learning rate |