Validity Without Ground Truth: What Stated-Preference Economics Offers the Evaluation of Language Models
Organizations: University of Michigan School of Information · Tisch University Professor, SC Johnson College of Business, Cornell University
Abstract
Many of the questions now put to large language models have no correct answer to score against: what a policy is worth, which option a user should choose, how to weigh competing values. Stated-preference economics has faced this problem for decades. It judges survey responses without knowing the true value, through a framework of validity and related concepts: content, construct, and criterion validity, reliability, incentive compatibility, and consequentiality. We argue that this framework is a general method for evaluating language models, and we set out what each concept means for LLM evaluation. We demonstrate the approach using a published water-quality stated preference economic valuation survey (Vossler et al. 2023) administered to six models. In this economic application, the validity tests take the form of predictions from economic theory: demand should slope down, and willingness to pay should respond to the scope of the good and to income. The tests separate the models sharply. Two older models fail the most basic test at a household income level of $75,000, and the two newest pass every test of theoretical validity we can score, but diverge on convergent validity. Passing validity tests shows that a model's answers are coherent, not that they are correct.
Figures & tables
| Concept | Question in stated preference | Question for LLM evaluation | In our demonstration |
|---|---|---|---|
| Content validity | Does the instrument describe the good, payment, and decision clearly and credibly? | Does the task pose the question it claims to, and could the model be recalling the benchmark rather than answering it? | Published survey text, verbatim; the scale’s published name withheld |
| Construct: theoretical | Do answers move as economic theory predicts? | Do answers respond to what should matter, in the right direction, and ignore what should not? | Demand slopes down; scope; income |
| Construct: convergent | Do different measures of the same value agree? | Do different prompts, formats, and scoring methods reach the same conclusion? | Two WTP estimators; referendum vs. open-ended; human estimates |
| Criterion validity | Do stated values match real payments? | Does what the model says match what it does when it acts? | Not assessed: the model never pays |
| Reliability | Does the answer survive repetition? | Are results stable across samples, reruns, and model versions? | Ten independent draws per question |
| Incentive compatibility | Does the format reward truthful answers? | Does the setup reward telling the user what they want to hear? | Referendum format; not tested |
| Local watershed | Non-local watershed | Full study region | |||||||
| Model | One-level | Min L2 | Min L3 | One-level | Min L2 | Min L3 | One-level | Min L2 | Min L3 |
| Claude Sonnet 5 | 1,155 | 1,808 | 1,220 | 462 | 651 | 485 | 1,463 | 1,753 | 1,058 |
| (90) | (114) | (100) | (53) | (51) | (54) | (119) | (105) | (68) | |
| GPT-5.6 Terra | 1,623 | 3,156 † | 1,184 | 925 | 801 | 541 | 2,066 | 3,156 † | 1,747 |
| (135) | (301) | (100) | (95) | (74) | (60) | (158) | (301) | (134) | |
| GPT-4o | 3,156 † | – | 2,945 | 936 | 957 | 617 | 2,524 | – | 2,657 |