Language models can notice an impossible engineering problem yet still report it as solved
Organizations: Santa Clara University Santa Clara, CA 95053, USA
Abstract
Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection of their flawed counterparts. Each reply required a "solved" or "cannot solve" status; rejection meant "cannot solve" or withholding an answer. The initial prompts did not warn that problems could be flawed. Across three recent models, 12 of 90 replies failed to reject a flawed problem. In 11 of these replies, the model stated the flaw, answered a corrected problem and still reported the original as "solved", according to artificial intelligence raters and numerical checks. We later retested four models from one provider, offering "flawed" instead of "cannot solve" and asking them to name and explain the defect. Three models showed statistically significant increases in rejection, but valid-problem solving fell in three. Evaluations therefore need to score both versions and distinguish flaw recognition from the reported status.
Figures & tables
| Family | Problem | What the two references do | Solve items | Flaw twin |
|---|---|---|---|---|
| T02 | Two-span beam, support settlement | Force method with kinematic theorems; stiffness method with linear programming of the static shakedown theorem | 5 | Load above collapse |
| T04 | Beam on a compression-only bed | Exact exponential basis with nested bracketing; transfer matrices and shooting | 5 | Gauge contradicts lift-off |
| T08 | Clamped cosine strip | Buckling-mode expansion, 2000 modes; Hermite elements, displacement control | 5 | Rise too small to snap |
| T11 | Turbine-disk spin test | Closed form with a root for the plastic boundary; Navier integration through the plastic zone | 5 | Overspeed above burst |
| T12 | Shrink-fitted disk on a shaft | Lamé fit plus the T11 disk; shooting for disk and shaft | 5 | Fit claimed tight is loose |
| Tier 1 | 13 single-topic families | Per family, two methods (Supplementary Table 1) | 52 | 20 premise pairs |
| Model | Provider | Harness | Tier 1 | Tier 2 solve | Flaws (30) | Repeats |
|---|---|---|---|---|---|---|
| Fable 5.1 | Anthropic | Claude CLI | complete | complete | complete | solve, flaws |
| Opus 5.5 | Anthropic | Claude CLI | complete | complete | complete | solve |
| Sonnet 5.5 | Anthropic | Claude CLI | complete | complete | complete | solve, flaws |
| Opus 5 | Anthropic | Claude CLI | complete | complete | complete | solve, flaws |
| Sonnet 5 | Anthropic | Claude CLI | complete | complete | complete | — |
| Opus 4.8 | Anthropic | Claude CLI | complete | complete | complete | — |
| Earlier unwarned | Fresh unwarned | Fresh schema only | ||||||||
| Model | first | repeats | rejected | consistent | joint | FA | rejected | consistent | joint | FA |
| Opus 5 | 23 | 26, 29 | 12/30 | 28 | 12 | 0 | 30/30 | 30 | 30 | 0 |
| Opus 4.8 | 7 | — | 9/30 | 24 | 8 | 0 | 30/30 | 19 | 19 | 2 |
| Sonnet 5.5 | 28 | 29, 28 | 27/30 | 29 | 26 | 1 | 30/30 | 25 | 25 | 0 |
| Sonnet 5 | 3 | — | 0/16 | 11 | 0 | 0 | 15/16 | 7 | 7 | 2 |
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Family | Problem |
|---|---|
| F01 | Thermal buckling versus first yield of a clamped stepped column |
| F02 | Two-span beam with a settling compression-only bearing and a thermal gradient |
| F03 | Torsional resonance of a three-disk rotor on a stepped shaft |
| F04 | Finite free–free beam on an elastic foundation, off-centre load |
| F05 | Fatigue versus first-cycle yield of a notched shaft, with a minimum-diameter redesign |
| F06 | Edge crack in a restrained strip under thermal contraction: fracture versus net-section yield |
| Claude Code CLI | Codex CLI 0.158.0 | Antigravity CLI 1.2.13 | |
|---|---|---|---|
| Models | Anthropic panel | GPT-6 and GPT-5 series | Gemini, gpt-oss-120b |
| System text | Replaces the system prompt | Prepended; the CLI keeps its own agent instructions | Prepended; the CLI keeps its own agent prompt |
| Tools | None | Disabled, read-only sandbox | Cannot be disabled; a call that logs a tool step is graded as a failure (nine Gemini 3.8 Flash calls) |
| Effort | High (Haiku 4.5: no flag) | High | Encoded in the model identifier |
| Output cap | Tier 2: 128,000 (Haiku 4.5: 64,000). Tier 1: CLI default | Service default | Not settable |
| Sampling | CLI default; no temperature or seed set | CLI default; no temperature or seed set | CLI default; no temperature or seed set |
| Model | Solve items | Unwarned rejection | False alarms |
|---|---|---|---|
| Fable 5.1 | 23/25 (75–98%) | 27/30 (74–97%) | 0/30 (0–11%) |
| Opus 5.5 | 23/25 (75–98%) | 30/30 (89–100%) | 0/30 (0–11%) |
| Sonnet 5.5 | 24/25 (80–99%) | 28/30 (79–98%) | 0/30 (0–11%) |
| Opus 5 | 25/25 (87–100%) | 23/30 (59–88%) | 0/30 (0–11%) |
| Sonnet 5 | 13/25 (33–70%) | 3/30 (3–26%) | 0/30 (0–11%) |
| Opus 4.8 | 19/25 (57–89%) | 7/30 (12–41%) | 0/30 (0–11%) |
| Outcome | Contrast | Only first | Only second | Holm |
|---|---|---|---|---|
| Solve items (exploratory) | Opus 5.5 vs Opus 5 | 0 | 2 | 1 |
| Solve items (exploratory) | Opus 5 vs Opus 4.8 | 6 | 0 | 0.16 |
| Solve items (exploratory) | Sonnet 5.5 vs Sonnet 5 | 11 | 0 | 0.0059 |
| Solve items (exploratory) | GPT-6-Sol vs GPT-5.6-Sol | 3 | 2 | 1 |
| Solve items (exploratory) | GPT-6-Astra vs GPT-6-Sol | 2 | 0 | 1 |
| Solve items (exploratory) | Opus 5.5 vs GPT-6-Astra | 0 | 2 | 1 |
| Replies | |
|---|---|
| Scored replies, primary panel runs | 3679 |
| Fenced JSON | 3643 |
| of which two JSON blocks | 1 |
| Bare JSON (all gpt-oss-120b) | 26 |
| Unparsed | 10 |
| Seeded sample, graded as stated | 50 of 50 |
| Scheme | Cases | Raw agreement | Cohen’s |
|---|---|---|---|
| Four classes (noted, misjudged, adopted, silent) | 232 | 0.88 | 0.81 |
| Three classes (noted, misjudged, unexamined) | 232 | 0.94 | 0.89 |