MechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative Physics
Organizations: Idiap Research Institute, Switzerland · École Polytechnique Fédérale de Lausanne (EPFL), Switzerland
Abstract
This work introduces MechReasoner, a mechanistic qualitative simulator grounded in confluence-based qualitative physics, together with a benchmark for mechanistic inference. Current large language models (LLMs) generate fluent mechanistic descriptions that do not reliably follow from underlying structural and causal constraints. The benchmark tests whether answers preserve simulator-licensed ambiguity, quantified claims, episode-graph transition evidence, repairs, and trace-support judgments. Its 1,120 items are generated deterministically from admissible interpretation sets, component states, scenario restrictions, confluence constraints, and derivation steps across 18 catalog mechanisms and six task families. Each mechanism undergoes converter checks of structure and topology and behavioral checks against quantitative simulations. GPT-5.5 accuracy decreases as family-specific mechanistic complexity increases, from 76.1% in the lowest-complexity bucket (B1) to 38.0% in the highest-complexity bucket (B4). The negative association remains after controls for rendered-prompt and expected-answer length. These results show that qualitative simulators can support auditable NLP benchmarks for mechanistic inference.
Figures & tables
| Task family | Decision criterion | Illustrative figure |
|---|---|---|
| State consistency | A candidate composite state is consistent iff there exists a complete qualitative interpretation satisfying the selected component states, operating context, case restrictions, and all active laws. | |
| Plausibility | A narrative is plausible (P) iff there exists an -step path from the initial episode such that the episode at every checkpoint satisfies all observations listed for that checkpoint; otherwise it is implausible (I). | P |
| Necessity | Every retained claim has at least one supplied reachable edge whose source satisfies all SOURCE facts. A claim is necessary (N) iff every such edge has a target satisfying at least one TARGET alternative; otherwise it is not necessary (U). | |
| Episode-graph transitions | From the reachable edges up to horizon , list every distinct component state change as a (component, source state, target state) triple. Exclude declared changes not realized by an edge and edges that change only the qualitative interpretation. | |
| Functional recovery | Classify each policy over all named post-repair outcomes and their one-step successors: unsafe (U) if any is unsafe; otherwise deadline failure (D) if any deadline episode fails the goal; otherwise recovered (R). | names one step policy outcome unsafe U |
| Trace faithfulness | Replay each trace in order. A step is unfaithful if its displayed before-set differs from the current set or if its cited law does not produce exactly its displayed after-set. Continue from the displayed after-set even after an unfaithful step, and count all such steps. | step 1 step 2 step 3 step 4 |
| Task family | Complexity scalar | What the scalar counts |
|---|---|---|
| State consistency | Active confluences and state and case restrictions for every displayed candidate–case pair. | |
| Plausibility | Displayed qualitative facts across the successor checkpoints in one candidate narrative. | |
| Necessity | Distinct accepted causal edges reachable within the task horizon. | |
| Episode-graph transitions | Possible boundary-cause and declared-transition pairings at each reachable source episode. | |
| Functional recovery | Displayed repair plans and their distinct accepted one-step effects. | |
| Trace faithfulness | Displayed trace-step occurrences across the independent traces. |
| Task family | Items | Mech. |
|---|---|---|
| State consistency | 212 | 17 |
| Plausibility | 136 | 11 |
| Necessity | 160 | 14 |
| Episode-graph transitions | 100 | 11 |
| Functional recovery | 288 | 18 |
| Trace faithfulness | 224 | 14 |
| Analysis-bin effect | Unadjusted | Adjusted |
|---|---|---|
| B2 vs B1 OR | 0.486 | 0.894 |
| B3 vs B1 OR | 0.259 | 0.734 |
| B4 vs B1 OR | 0.073 | 0.201 |
| Ordinal coefficient | -0.822 | -0.479 |
| Ordinal odds ratio | 0.439 | 0.619 |
| 95% OR interval | [0.370, 0.522] | [0.547, 0.701] |
| Task family | Majority/constant | GPT-5.5 | |||||
|---|---|---|---|---|---|---|---|
| Accuracy | Partial | E2E accuracy [95% CI] | E2E partial [95% CI] | Semantic accuracy | Prot. | Infra. | |
| Necessity | 0.0 | 50.0 | 23.1 [12.8, 34.8] | 65.6 [59.5, 72.0] | 23.1 | 0 | 0 |
| Plausibility | 0.0 | 50.0 | 36.0 [26.8, 43.8] | 89.2 [86.0, 92.0] | 37.7 | 6 | 0 |
| State consistency | 0.0 | 67.5 | 79.7 [71.9, 87.3] | 98.4 [97.0, 99.3] | 80.1 | 1 | 0 |
| Episode-graph transition | 0.0 | 0.0 | 78.0 [66.2, 94.0] | 88.3 [82.4, 96.6] | 78.0 | 0 | 0 |
| Trace faithfulness | 4.5 | 4.5 | 75.4 [70.1, 81.2] | 75.4 [70.1, 81.2] | 75.4 | 0 | 0 |
| Model | SC | P | N | T | FR | TF | Prot. |
|---|---|---|---|---|---|---|---|
| GPT-5.5 | 79.7 | 36.0 | 23.1 | 78.0 | 62.5 | 75.4 | 8 |
| o3 | 25.9 | 5.1 | 6.2 | 57.0 | 46.9 | 43.3 | 52 |
| DeepSeek-R1 | 7.1 | 0.0 | 6.2 | 45.0 | 39.6 | 2.2 | 110 |
| gpt-oss-20b | 0.0 | 0.0 | 1.9 | 39.0 | 29.5 | 0.0 | 152 |
| Qwen3-Coder-30B | 0.9 | 0.0 | 18.1 | 5.0 | 10.4 | 0.0 | 107 |
| GPT-4.1 | 0.5 | 0.0 | 3.1 | 15.0 | 12.8 | 0.0 | 24 |
| Family | Candidate construction | Semantic admission and fixed controls |
|---|---|---|
| State consistency | Complete composite-state rows are drawn from the mechanism state space and paired with persistent context plus scenario-specific restrictions. Inconsistent rows are deliberately not prefiltered. | The exact intrastate relation is queried for every candidate–scenario pair. Reachability and transition rules are outside this family. Retained candidates must fall in the frozen absolute D1–D4 constraint-count bands. |
| Plausibility | One complete initial episode and its exact bounded graph yield 32 checkpoint narratives. Positive packets come from one complete path and layer invariants; negative packets are constructed against exact non-witness certificates. | Every narrative contains three to six facts at every checkpoint, with 16 and 16 labels. All facts in a positive narrative share one exact-length witness. Horizons are 1, 1, 2, and 3 for D1–D4. |
| Necessity | Four claims are mined over one exact reachable edge set. Antecedents contain three or four source facts and consequents contain two or three distinct target alternatives. | Antecedents are nonvacuous. Necessary claims require the joint disjunction and reject a sufficient proper disjunction; unnecessary claims retain an exact counterexample, including one at the deepest expanded source layer. The normal task has a two–two balance. |
| Episode-graph transitions | A complete initial episode and horizon determine an exact bounded graph; each candidate answer projects distinct component-state changes witnessed by its edges. Seeded surface variants change evidence order only. | The answer must be nonempty. Unrealized declarations and interpretation-only edges are excluded, while simultaneous component changes are projected individually. Preferred horizons are 1, 1, 2, and 2 for D1–D4. |
| Functional recovery | Removing the declared initial-coordinate restrictions produces the exhaustive fault-belief episode library for the task. Externally observable goal and unsafe atoms define contracts, and each policy names its complete repaired outcome set. | Labels quantify over every named outcome and every exact one-edge successor. D1–D4 display 1, 3, 5, and 15 plans; D4 contains five each of , and policy order is seeded before taking the level prefix. |
| Trace faithfulness | Exact local confluence support generates six-step sequential traces. Faithful steps use the exact before-domain and supported after-set; unfaithful steps apply a controlled wrong law, before-domain mismatch, or near-miss supported set and are rechecked by the same exact solver. | Trace scenarios are independent but steps within a trace update sequentially. D1–D4 contain 22, 29, 37, and 44 traces, hence 132, 174, 222, and 264 displayed step judgments. |
| Model | Scale | Architecture | Documented profile | Capability class |
|---|---|---|---|---|
| GPT-5.5 ( OpenAI 2026 ) | Frontier (undisclosed) | Undisclosed GPT architecture | Reasoning, coding, tool-heavy agents, and long-running tasks; 1.05M context | Agentic-capable reasoning LLM |
| o3 ( OpenAI 2025d ) | (undisclosed) | Undisclosed reasoning architecture | Multi-step reasoning across text, code, and images; function calling | Agentic-capable reasoning LLM |
| DeepSeek-R1 ( DeepSeek-AI 2025 ) | 671B total / 37B active | Sparse MoE transformer | Reasoning model based on DeepSeek-V3-Base with cold-start data and reinforcement learning | Reasoning LLM |
| gpt-oss-20b ( OpenAI 2025b ) | 21B total / 3.6B active | Sparse MoE transformer | Reasoning post-training and interleaved web, Python, and developer-tool use | Agentic-capable reasoning LLM |
| Qwen3-Coder-30B ( Qwen Team 2025b ) | 30.5B total / 3.3B active | Sparse MoE transformer | Code-specialized, repository-scale, agentic coding and browser use | Agentic-capable coding LLM |
| GPT-4.1 ( OpenAI 2025c ) | (undisclosed) | Undisclosed GPT architecture | Instruction following, coding, tool use, long context, and powering agents | Agentic-capable LLM |
| Display label | Exact evaluator profile | Effective output-token limit | Reasoning-effort request |
|---|---|---|---|
| GPT-5.5 | gpt-5.5 | 20,000 | low |
| o3 | o3 | 20,000 | low |
| DeepSeek-R1 | DeepSeek-R1 | 20,000 | not sent |
| gpt-oss-20b | openai/gpt-oss-20b | 20,000 | low |
| Qwen3-Coder-30B | qwen/qwen3-coder-30b-a3b-instruct | 20,000 | not sent |
| GPT-4.1 | gpt-4.1 | 20,000 | not sent |
| Catalog mechanism id | Provenance | Supported task families | Retained |
|---|---|---|---|
| electrical_electrothermal_resistor_parallel | Repository composition | SC, FR, TF | 48 |
| electrical_resistor | MSL 4.1.0 | SC, P, N, T, FR, TF | 68 |
| electrical_electrothermal_resistor_ladder | Repository composition | SC, P, N, FR, TF | 69 |
| mechanics_branched_three_mass | Repository composition | SC, FR, TF | 44 |
| mechanics_compare_braking_force | MSL 4.1.0 | SC, P, N, T, FR, TF | 96 |
| mechanics_compare_braking_torque | MSL 4.1.0 | SC, P, N, T, FR, TF | 96 |