Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection of their flawed counterparts. Each reply required a "solved" or "cannot solve" status; rejection meant "cannot solve" or withholding an answer. The initial prompts did not warn that problems could be flawed. Across three recent models, 12 of 90 replies failed to reject a flawed problem. In 11 of these replies, the model stated the flaw, answered a corrected problem and still reported the original as "solved", according to artificial intelligence raters and numerical checks. We later retested four models from one provider, offering "flawed" instead of "cannot solve" and asking them to name and explain the defect. Three models showed statistically significant increases in rejection, but valid-problem solving fell in three. Evaluations therefore need to score both versions and distinguish flaw recognition from the reported status.
Figures & tables
Figure 1: What a reply to a flawed twin does. a , One evaluation pair from the no-tension-bed family (T04). The statements differ only in the dial-gauge datum. The sketch draws the pair to scale along the beam: the left 2.15 m of the beam lifts off the compression-only bed, so a gauge reading of −3.49 mm (upward) agrees with the mechanics and 2.84 mm downward contradicts them. The other data fix a lift-off of the left end, which OpenSees (Open System for Earthquake Engineering Simulation) reproduces. In the consistent twin the gauge value equals the graded uplift of that end, so that twin states one graded answer (Methods). b , Four replies to the flawed twin, unwarned, one per outcome class, with what the JavaScript Object Notation (JSON) answer contains. Quotes are verbatim, with typography normalised. GPT-6-Astra rejects the statement. Fable 5.1 names the inconsistency, returns the lift-off solution with the gauge set aside, and marks the problem solved. Sonnet 5 treats the gauge as confirmation that the end stayed down. Haiku 4.5 builds its answer on the gauge. c , Outcome of every unwarned flawed twin, Tier 2 (30, multi-regime) and Tier 1 (20, standard single-topic): rejected; noted, split on Tier 2 into a repaired input and answers for the flawed state; checked with a wrong conclusion; unexamined; or no usable answer. Classes and the split are from blinded coding ( κ=0.89 and 0.92; Methods).
Figure 2: Benchmark architecture. a , Parametric generators produce five multi-regime and 13 single-topic families. Two independent reference methods and the acceptance gates check each key; OpenSees provides 62 third-party checks for the two-span beam (T02), no-tension bed (T04) and clamped cosine strip (T08). Each evaluation set was frozen before its model calls. A twin pair changes one datum or premise between a consistent and a flawed statement. b , The 14-model panel receives solve, critique and premise-check prompts. A deterministic grader reads the final structured answer, applies the numerical tolerance and scores rejection. The first scored reply counts; harness settings and prompt variants are specified in Methods. c , Separate analyses measure solve accuracy, rejection, recognition and repair. Blinded AI coding describes the misses; paired tests compare models on the same items and on pairs whose consistent twins both models solved. Repeats measure consistency across three runs. CLI, command line interface; JSON, JavaScript Object Notation.
Figure 3: The five Tier-2 families. a , Sketches. T02, two-span steel beam with a settled middle support and plastic hinges. T04, free–free beam on a bed that pushes and never pulls; in the instance drawn, the left end lifts. T08, clamped cosine strip under a midpoint force; the dashed curve is the inverted stable state. T11, annular turbine disk in a spin test, plastic zone shaded, blade load at the rim. T12, disk shrink-fitted on a shaft. b , Reference-A solutions for one development instance per family, with the answers of the textbook shortcuts. Each shortcut misses at least one graded answer by more than 10%; in T08 the one-mode shortcut lands within 0.3% of the critical force and misses the clamp moment by 68%. Shaded bands mark the plastic-moment range (T02) and the contact zone (T04). c , One flaw twin per family. Only the stated datum or premise differs from the consistent twin. The five instances drawn here are development items, not evaluation items. Six evaluation pairs are used per family (30 pairs).
Family
Problem
What the two references do
Solve items
Flaw twin
T02
Two-span beam, support settlement
Force method with kinematic theorems; stiffness method with linear programming of the static shakedown theorem
5
Load above collapse
T04
Beam on a compression-only bed
Exact exponential basis with nested bracketing; transfer matrices and shooting
5
Gauge contradicts lift-off
T08
Clamped cosine strip
Buckling-mode expansion, 2000 modes; Hermite elements, displacement control
5
Rise too small to snap
T11
Turbine-disk spin test
Closed form with a root for the plastic boundary; Navier integration through the plastic zone
5
Overspeed above burst
T12
Shrink-fitted disk on a shaft
Lamé fit plus the T11 disk; shooting for disk and shaft
5
Fit claimed tight is loose
Tier 1
13 single-topic families
Per family, two methods (Supplementary Table 1)
52
20 premise pairs
Table 1: Benchmark composition and verification. Two independent methods verify each key. Acceptance gates test decision margins, shortcut errors, alternative readings and numerical sensitivity (Methods). Supplementary Table 8 lists the graded quantities. T11 and T12 have no OpenSees check.
Model
Provider
Harness
Tier 1
Tier 2 solve
Flaws (30)
Repeats
Fable 5.1
Anthropic
Claude CLI
complete
complete
complete
solve, flaws
Opus 5.5
Anthropic
Claude CLI
complete
complete
complete
solve
Sonnet 5.5
Anthropic
Claude CLI
complete
complete
complete
solve, flaws
Opus 5
Anthropic
Claude CLI
complete
complete
complete
solve, flaws
Sonnet 5
Anthropic
Claude CLI
complete
complete
complete
—
Opus 4.8
Anthropic
Claude CLI
complete
complete
complete
—
Table 2: Model panel. Time limit per call: 1800 s on Tier 1 and 2700 s on Tier 2; the first scored reply per item counts (Methods). Tools, web search and user configuration are off where the harness allows it (Supplementary Table 2). Model identifiers are in Supplementary Table 9, output caps in Supplementary Table 2, and the handling of calls that ended without a reply in the Methods. Repeats are three independent runs. CLI, command line interface.
Figure 4: Standard problems bunch the panel; multi-regime problems spread it. a , Solve accuracy on Tier 1 (open circles, 52 items) and Tier 2 (filled circles, 25 items). Bars are Wilson 95% intervals. Model names are coloured by provider. Eleven of the 13 models run on Tier 1 solve 45 to 52 of 52; Haiku 4.5 (13) and gpt-oss-120b (11) do not. Gemini 3.1 Pro was not run on Tier 1. On Tier 2 the same panel runs from 0 to 25 correct (Cochran’s Q=173 , 13 degrees of freedom, P=4.6×10−30 , 14 models). b , What each Tier-2 answer is: correct; near-miss (every quantity within 5%); a named textbook shortcut; another error; or no usable answer (time limit, output budget or harness timeout).
Figure 5: Unwarned rejection in the first collection. a , Share of flawed twins rejected (the reply marks the statement unsolvable or withholds an answer). Filled circles: Tier 2, unwarned, 30 twins, with Wilson 95% intervals; counts at the right. Ticks: Tier 1 premise twins, unwarned, 20 pairs. b , False alarms: consistent Tier-2 twins rejected with no warning, out of 30. c , The warned arm on the original nine pairs: rejections of the same nine flawed twins without (open circle) and with (diamond) the warning; orange numbers are the consistent twins rejected with the warning. The warning changes both the instruction and the status field; Supplementary Table 15 separates them for five models. d , Six contrasts fixed before the 21 added pairs existed. Bars count discordant items (only the first model right, right of zero; only the second, left). Grey: solve accuracy, 25 items (exploratory). Blue: rejection on the 21 added pairs (confirmatory). Hatched: the same 21 pairs restricted to those whose consistent twin both models solved (exploratory). P values are exact McNemar tests (two-sided binomial tests on the discordant items), Holm-adjusted for multiple testing within each set of six; values below 0.001 are in Supplementary Table 4.
Figure 6: Reliability and time. a , Items solved (25 solve items) or flawed twins rejected (30; nine for Gemini 3.8 Flash) in every one of three runs ( pass3 , squares) against the individual runs (circles). Solved at least once in three runs (pass@3): Fable 5.1, Opus 5.5, Opus 5 and GPT-6-Astra 25 of 25, Sonnet 5.5 and GPT-6-Sol 24, Gemini 3.8 Flash 19. Item-level repeats are in Supplementary Fig. 6. b , Tier-2 solve accuracy against median wall time per item. c , The same solve accuracy against median output tokens as each harness defines them. Token definitions differ by harness (Supplementary Table 2), so wall time is the common unit. Across this panel, at one effort setting per model, accuracy is not ordered by time.
Earlier unwarned
Fresh unwarned
Fresh schema only
Model
first
repeats
rejected
consistent
joint
FA
rejected
consistent
joint
FA
Opus 5
23
26, 29
12/30
28
12
0
30/30
30
30
0
Opus 4.8
7
—
9/30
24
8
0
30/30
19
19
2
Sonnet 5.5
28
29, 28
27/30
29
26
1
30/30
25
25
0
Sonnet 5
3
—
0/16
11
0
0
15/16
7
7
2
Table 3: Second collection of the two confirmatory contrasts. Unwarned prompt: the reply can report “solved” or “cannot solve”. Schema-only prompt: “solved” or “flawed” with a defect type, and no check instruction. Earlier runs: 29–30 September, flawed twins rejected of 30. Fresh runs: times are in Coordinated Universal Time (UTC); calls started 30 September 21:31 UTC to 1 October 07:43 UTC, both arms of each model started together, compared on the same pairs (30, or for Sonnet 5 the 16 pairs with a final unwarned call). Rejected: flawed twins rejected. Consistent: consistent twins solved. Joint: consistent twin solved and flawed twin rejected. FA: false alarms on consistent twins. Over all 29 of its schema-only pairs Sonnet 5 rejected 25, solved 12 consistent twins and raised 3 false alarms. Prespecified tests are in Supplementary Table 16.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Supplementary Figure 1: Tolerance sensitivity. Solve items counted correct when every key tolerance is multiplied by a factor from 0.25 to 10. The frozen rule is the factor 1. Colour is the change from that rule. No model pair reverses order between factors 1 and 5. The largest move in that range is four additional items, for Sonnet 5.
Supplementary Figure 2: Third-party verification. Relative difference between OpenSees and the answer key for every checked quantity, with counts of the qualitative checks (shakedown bracket, flaw reproduced, mode match). T08 compares exact beam kinematics with the keys’ shallow-arch model. The two points above 1% are imperfection-sensitive bifurcation loads, accepted at a 3% tolerance. Sixty-two checks, no failures. T11 and T12 are not in this figure.
Supplementary Figure 3: Tier-2 solve accuracy by family. Items correct out of five in each family. On the primary run, Gemini 3.8 Flash solved none of the five T08 items.
Supplementary Figure 4: Tier-1 flaw handling. Critique pairs (detect the implanted error, locate the first wrong step, false alarm on the clean solution) and premise pairs, unwarned and warned (reject: the flawed statement is rejected; false alarm: the consistent statement is rejected). Premise rejection shows the same warned-versus-unwarned gap as Tier 2, at 20 pairs.
Supplementary Figure 5: Coder agreement. Confusion matrix of the two blinded coders over 232 missed flaws. Three-class κ=0.89 . Four-class κ=0.81 . The off-diagonal mass in the four-class matrix sits on the adopted/silent boundary, which the primary scheme merges.
Supplementary Figure 6: Item-level consistency. Each Tier-2 solve item in each of three runs, for the seven models that were repeated. A cell is one item in one run. GPT-6-Astra has no flips. GPT-6-Sol flips on eight items, which is why pass3 falls from a primary score of 23/25 to 16/25.
Family
Problem
F01
Thermal buckling versus first yield of a clamped stepped column
F02
Two-span beam with a settling compression-only bearing and a thermal gradient
F03
Torsional resonance of a three-disk rotor on a stepped shaft
F04
Finite free–free beam on an elastic foundation, off-centre load
F05
Fatigue versus first-cycle yield of a notched shaft, with a minimum-diameter redesign
F06
Edge crack in a restrained strip under thermal contraction: fracture versus net-section yield
Appendix
Supplementary Table 1: Tier-1 families. Hard-10 is the hardest tenth of a 1280-item pool, selected by a model-free score and frozen before any model call. The evaluation slice used here is 52 solve items, 21 critique pairs and 20 premise pairs.
Claude Code CLI
Codex CLI 0.158.0
Antigravity CLI 1.2.13
Models
Anthropic panel
GPT-6 and GPT-5 series
Gemini, gpt-oss-120b
System text
Replaces the system prompt
Prepended; the CLI keeps its own agent instructions
Prepended; the CLI keeps its own agent prompt
Tools
None
Disabled, read-only sandbox
Cannot be disabled; a call that logs a tool step is graded as a failure (nine Gemini 3.8 Flash calls)
Supplementary Table 2: Harness settings that affect comparability. Wall time is the common unit across providers. Token counts are not. “Output tokens” include thinking in all three harnesses, on different tokenizers. CLI, command line interface.
Model
Solve items
Unwarned rejection
False alarms
Fable 5.1
23/25 (75–98%)
27/30 (74–97%)
0/30 (0–11%)
Opus 5.5
23/25 (75–98%)
30/30 (89–100%)
0/30 (0–11%)
Sonnet 5.5
24/25 (80–99%)
28/30 (79–98%)
0/30 (0–11%)
Opus 5
25/25 (87–100%)
23/30 (59–88%)
0/30 (0–11%)
Sonnet 5
13/25 (33–70%)
3/30 (3–26%)
0/30 (0–11%)
Opus 4.8
19/25 (57–89%)
7/30 (12–41%)
0/30 (0–11%)
Appendix
Supplementary Table 3: Tier-2 rates with Wilson 95% intervals. Solve items are the 25 Tier-2 solve problems. Unwarned rejection (flawed twins) and false alarms (consistent twins) are out of 30 pairs for all 14 models. Gemini 3.1 Pro was not run on Tier 1 and is included here.
Outcome
Contrast
Only first
Only second
Holm P
Solve items (exploratory)
Opus 5.5 vs Opus 5
0
2
1
Solve items (exploratory)
Opus 5 vs Opus 4.8
6
0
0.16
Solve items (exploratory)
Sonnet 5.5 vs Sonnet 5
11
0
0.0059
Solve items (exploratory)
GPT-6-Sol vs GPT-5.6-Sol
3
2
1
Solve items (exploratory)
GPT-6-Astra vs GPT-6-Sol
2
0
1
Solve items (exploratory)
Opus 5.5 vs GPT-6-Astra
0
2
1
Appendix
Supplementary Table 4: Contrasts. Exact McNemar tests, Holm-adjusted within the six contrasts for that outcome. Counts are discordant items (only the first model right, then only the second). The six contrasts were fixed on 29 September 2026, after the solve items and the original nine flaw pairs had been scored and before the 21 added pairs existed. Solve accuracy is exploratory. Rejection on the 21 added pairs is the confirmatory test. Rejection on all 30 pairs is the combined estimate. The same contrasts under the twin controls are in Supplementary Table 14.
Replies
Scored replies, primary panel runs
3679
Fenced JSON
3643
of which two JSON blocks
1
Bare JSON (all gpt-oss-120b)
26
Unparsed
10
Seeded sample, graded as stated
50 of 50
Appendix
Supplementary Table 5: Parser census, scored replies. The grader uses the last fenced JavaScript Object Notation (JSON) block and also accepts bare JSON. A seeded random sample of 50 replies, stratified by harness (seed 20260930), matched the graded value in every case. The flagged set (0 unit-scale slips, 4 sign flips, 35 answers whose correct value appears elsewhere in the reply) was reviewed separately; none was a parser error. The census covers the primary panel runs only. The single-part arms (180 calls) gave 178 fenced replies and two non-answers; the second collection (456 final calls) gave 447 fenced replies, one unparsed reply and eight non-answers.
Scheme
Cases
Raw agreement
Cohen’s κ
Four classes (noted, misjudged, adopted, silent)
232
0.88
0.81
Three classes (noted, misjudged, unexamined)
232
0.94
0.89
Appendix
Supplementary Table 6: Coding of missed flaws. Both coders labelled 232 cases. The three-class scheme merges adopted and silent into “unexamined” and was adopted after the four-class disagreements were seen. The 28 four-class disagreements were adjudicated by GPT-6-Astra, neither a coder nor the coding assistant, blind to model, run, item and the study’s aims. A first adjudication by Claude Opus 5.5, the model behind the coding assistant, gave the same class on 26 of 28 cases and on 28 of 28 under the three-class scheme. An earlier unblinded labelling of 42 cases by the coding assistant, made before the codebook was fixed, is not used; it agrees with the two coders at three-class κ=0.76 and 0.69 (four-class 0.39 and 0.37), and all 13 cases it called noted are noted for both coders.