cs.AISep 28, 2026

MechReasoner: A Simulator and Benchmark for Mechanistic Reasoning in Qualitative Physics

Authors: Danilo Gusicuma, André Freitas

Organizations: Idiap Research Institute, Switzerland · École Polytechnique Fédérale de Lausanne (EPFL), Switzerland

Abstract

This work introduces MechReasoner, a mechanistic qualitative simulator grounded in confluence-based qualitative physics, together with a benchmark for mechanistic inference. Current large language models (LLMs) generate fluent mechanistic descriptions that do not reliably follow from underlying structural and causal constraints. The benchmark tests whether answers preserve simulator-licensed ambiguity, quantified claims, episode-graph transition evidence, repairs, and trace-support judgments. Its 1,120 items are generated deterministically from admissible interpretation sets, component states, scenario restrictions, confluence constraints, and derivation steps across 18 catalog mechanisms and six task families. Each mechanism undergoes converter checks of structure and topology and behavioral checks against quantitative simulations. GPT-5.5 accuracy decreases as family-specific mechanistic complexity increases, from 76.1% in the lowest-complexity bucket (B1) to 38.0% in the highest-complexity bucket (B4). The negative association remains after controls for rendered-prompt and expected-answer length. These results show that qualitative simulators can support auditable NLP benchmarks for mechanistic inference.

Figures & tables

Explore similar work

CardsList
  1. Simulate, Reason, Decide: Scientific Reasoning with LLMs for Simulation-Driven Decision Making

    Jun 3, 2026Yuhan Yang, Ruipu Li, Alexander RodríguezNeuro-Symbolic FrameworkDecisions

  2. MechBench: Can AI Scientific Agents Discover Mechanisms Beyond Phenomenal Laws?

    Sep 28, 2026Zihan Yu, Jiadong Zhang, Jialin Cheng +2Scientific Agents

  3. DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

    May 25, 2026Matt L. Wiemann, Lindsay M. Smith, Peter Melchior +4Large Language Model BenchmarksLarge Language Model Agents