Investigating Human--AI Discrepancies via Multiple-Solution Problems
Organizations: Department of Mathematics, Stanford University · Institute of Computational and Mathematical Engineering, Stanford University · Department of Statistics and Department of Mathematics, Stanford University
Abstract
Frontier artificial intelligence (AI) models are benchmarked on whether they reach a correct answer. Yet many problems admit several correct answers and repeated attempts, by different people or by the same model resampled, trace out a distribution over them. In this work, we ask whether human and model reasoning lead to different distributions over valid solutions. Our testbed comprises 270 reasoning puzzles across five puzzle families. These multiple-solution puzzles each have 3 to 8 valid solutions and are simple enough that humans and models can solve them reliably. The resulting distributions differ markedly: models differ from one another, yet resemble each other far more than they resemble humans. Model distributions are, moreover, within every puzzle family, less diverse than human ones. We compare these discrepancies across puzzle categories, and trace how they respond to reasoning-effort settings, to prompting, and to perturbations of the puzzle that leave its solutions unchanged. Together, these results point at significant differences between human and AI problem-solving processes, and their choice among equally defensible solutions. As progressive deployment of AI systems in society comes into focus, evaluating such differences (beyond one-dimensional accuracy metrics) is increasingly important. Data and code are available at https://hai-discrepancies.github.io/
Figures & tables
| Mean normalized entropy [95% CI] | ||||
|---|---|---|---|---|
| Puzzle family | Human | GPT | Claude | Gemini |
| Arithmetic | ||||
| Maze | ||||
| Rooks | ||||
| Minesweeper | ||||
| Sudoku | ||||
| Arithmetic | |
| task | Make exactly 24 using the four displayed numbers. |
| rules | Use each number exactly once. Allowed operations: +, -, *, /. Parentheses are allowed. Do not use other numbers. Only one correct answer is needed. |
| output | Return exactly one line: ANSWER: expression=<expression>. Replace placeholders with your answer. |
| Maze | |
| task | Find one shortest path from S to G. |
| rules | Move only U/D/L/R. White cells are open; black cells are walls. S and G are open. The path must be shortest. Only one correct answer is needed. |
| Two puzzles: Minesweeper | |
| task | Solve Puzzle A and Puzzle B. |
| rules | Use the rule statement shown above each Minesweeper board. Only one correct answer is needed for each puzzle. |
| output | Return exactly two lines and nothing else. Replace placeholders with your answers. |
| line 1 | A: cell=<row,column> |
| line 2 | B: cell=<row,column> |
| Above each board, under its Puzzle A / Puzzle B label: | |
| Plain prompt, single-puzzle trials | |
| Solve the puzzle shown in the image. | |
| Return exactly one final answer line and nothing else. | |
| Your final line must exactly match the answer format shown in the image, including labels, punctuation, separators, and capitalization. | |
| Replace placeholders with your answer. | |
| Do not include any explanation, reasoning, preface, markdown, extra whitespace, or additional lines. | |
| Give the first correct answer that comes to mind. | |
| GPT | Claude | Gemini | |
| Model identifier | gpt-5.6-sol | claude-opus-4-8 | gemini-3.5-flash |
| API | Responses | Messages | GenerateContent |
| Reasoning control | reasoning effort | adaptive thinking | thinking level |
| Settings | low, medium | low, medium | low, medium |
| Output token cap | 8,192 | 8,192 | 8,192 |
| Text verbosity | low | — | — |
| Source | Responses | Correct/valid | Collected |
|---|---|---|---|
| Human, main battery (104 participants) | 10,400 | 90.5% | 2026-06 |
| Human, modules (417 participants, 2 waves) | 25,647 | 92.3% | 2026-06 |
| GPT, main battery | 40,000 | 98.0–99.4% | 2026-07 |
| GPT, modules | 68,000 | 99.8–99.9% | 2026-07 |
| Claude, main battery | 40,000 | 99.6–99.8% | 2026-06 |
| Claude, modules | 68,000 | 99.3–99.4% | 2026-06 |
| 95% bootstrap CI | ||||
|---|---|---|---|---|
| Comparison | Mean TV | Puzzles | Trials | Both |
| Uniform vs Human | ||||
| Uniform vs GPT | ||||
| Uniform vs Claude | ||||
| Uniform vs Gemini | ||||
| Human vs GPT | ||||
| Pearson correlation [95% CI] | ||||||
|---|---|---|---|---|---|---|
| Source pair | Mean | Arithmetic | Maze | Rooks | Minesweeper | Sudoku |
| Human vs GPT | ||||||
| Human vs Claude | ||||||
| Human vs Gemini | ||||||
| GPT vs Claude | ||||||
| GPT vs Gemini | ||||||
| Contrast | Mean | Arithmetic | Maze | Rooks | Minesweeper | Sudoku |
|---|---|---|---|---|---|---|
| GPT | ||||||
| Persona plain (low effort) | [1pt] | [1pt] | [1pt] | [1pt] | [1pt] | [1pt] |
| Persona plain (medium effort) | [1pt] | [1pt] | [1pt] | [1pt] | [1pt] | [1pt] |
| Medium low (plain prompt) | [1pt] | [1pt] | [1pt] | [1pt] | [1pt] | [1pt] |
| Medium low (persona prompt) | [1pt] | [1pt] | [1pt] | [1pt] | [1pt] | [1pt] |
| Claude | ||||||
| a Selected-solution prediction | ||||
|---|---|---|---|---|
| Family | Human | GPT | Claude | Gemini |
| Arithmetic | ||||
| Maze | ||||
| Rooks | ||||
| Minesweeper | ||||
| Sudoku | ||||
| Source | Acc. | |||
|---|---|---|---|---|
| Human | ||||
| GPT | ||||
| Claude | ||||
| Gemini |
| Model/Human | BA | |||
|---|---|---|---|---|
| GPT | ||||
| Claude | ||||
| Gemini |
| Source | Acc. | |||
|---|---|---|---|---|
| Human | ||||
| GPT | ||||
| Claude | ||||
| Gemini |
| Model/Human | BA | |||
|---|---|---|---|---|
| GPT | ||||
| Claude | ||||
| Gemini |
| Source | Acc. | |||
|---|---|---|---|---|
| Human | ||||
| GPT | ||||
| Claude | ||||
| Gemini |
| Model/Human | BA | |||
|---|---|---|---|---|
| GPT | ||||
| Claude | ||||
| Gemini |
| Source | Acc. | ||
|---|---|---|---|
| Human | |||
| GPT | |||
| Claude | |||
| Gemini |
| Model/Human | BA | ||
|---|---|---|---|
| GPT | |||
| Claude | |||
| Gemini |
| Source | Acc. | ||
|---|---|---|---|
| Human | |||
| GPT | |||
| Claude | |||
| Gemini |
| Model/Human | BA | ||
|---|---|---|---|
| GPT | |||
| Claude | |||
| Gemini |
| Module | Human | GPT | Claude | Gemini |
|---|---|---|---|---|
| Highlighting | 0.0917 (10/20) | 0.3034 (10/20) | 0.0146 (13/20) | 0.0935 (13/20) |
| Related context | 0.0010 (9/10) | 0.8383 (1/10) | 0.3363 (3/10) | 0.1878 (2/10) |
| Strategy primer | 0.2549 (0/10) | 0.6074 (2/10) | 0.3044 (3/10) | 0.2272 (3/10) |
| Spatial reflection | 0.1103 (6/30) | 0.0337 (28/30) | 0.0133 (28/30) | 0.0255 (28/30) |
| Number ordering | 0.2332 (1/10) | 0.1525 (5/10) | 0.0424 (4/10) | 0.1556 (2/10) |
| Related | Unrelated | |||
|---|---|---|---|---|
| Source | Mean | Mean | ||
| Human | 0.0086 | 9/10 | 0.0083 | 10/10 |
| GPT | 0.8381 | 1/10 | 0.8002 | 2/10 |
| Claude | 0.3297 | 6/10 | 0.5391 | 2/10 |
| Gemini | 0.2254 | 6/10 | 0.3830 | 5/10 |