Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy
Organizations: Harvard T.H. Chan School of Public Health
Abstract
When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state. We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 jurisdictions) and three exactly defined Medicaid income-eligibility quantities. Gold values come from an official data book and agree with an independent source in 101 of 102 checked cells. Under a pre-registered protocol, Claude Sonnet 5.5 and GPT-5.6 Sol reproducibly give another state's current value, identical across two independent repeats, for 10 and 25 of 153 items. Attribution is fragile, however. Crediting any wrong answer that equals another state's value yields 3-5x more reproducible substitutions than checking every number in the asked state's own records, because many apparent cross-state answers are the asked state's own values under another convention or from an earlier year. Claims about cross-jurisdiction error need a complete same-state reference set. We will release the protocol, gold table, and all model outputs.
Figures & tables
| Acc. [95% CI] | Alt | Stale | Anch | Oth | Uns | Unst | ||
|---|---|---|---|---|---|---|---|---|
| Sonnet | Preg. | .59 [.49,.68] | 1 | 6 | 2 | 4 | 0 | 11 |
| Infant | .48 [.39,.58] | 1 | 4 | 0 | 3 | 9 | 15 | |
| Parent | .21 [.14,.29] | 9 | 12 | 6 | 3 | 2 | 9 | |
| GPT | Preg. | .57 [.47,.66] | 0 | 2 | 0 | 12 | 6 | 3 |
| Infant | .68 [.58,.76] | 0 | 1 | 0 | 9 | 4 | 4 | |
| Parent | .28 [.20,.37] | 6 | 13 | 6 | 4 | 5 | 4 |