Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI
Organizations: Siemens AG · Technical University of Munich
Abstract
Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiomatised OWL 2 ontology, with correct answers grounded in the ontology by design. Distractors are generated by perturbing the right-hand-side class expressions of class definition axioms, and their incorrectness is formally verified by an OWL reasoner via entailment checks. We evaluate the pipeline on three ontologies: Pizza (small, academic), PMDco (complex, materials science), and DOID (large, biomedical), generating 112, 2,491, and 15,216 MCQs respectively. Distractors span four semantic categories from class unsatisfiability to weakened subsumptions, enabling diagnostic evaluation of specific reasoning failures. Items meet natural language quality standards: mean LLM judge scores of 4.02, 4.36, and 3.36 out of 5 confirm fluency, and correct-answer-to-distractor similarity above 0.8 shows that wrong options cannot be dismissed on surface form alone. Six LLMs evaluated zero-shot achieve 41.1-76.8% accuracy, well above the 25% random-guessing baseline, indicating the benchmarks are challenging and discriminative. This work is a step towards more reliable benchmarks for assessing logical reasoning in scientific AI.
Figures & tables
| Op | Transformation | Source Set | Guard / Cap |
|---|---|---|---|
| Per distinct filler at position : for each | Siblings of and their descendants to depth in | Skip if contains ; depth cap | |
| Per distinct filler at position : for each | All ancestors of in ; all levels | None | |
| Per distinct filler at position : for each | Direct children of in | Cap children per filler | |
| Per distinct filler at position : for each | Classes disjoint with in | Skip if contains ; cap per filler | |
| = cardinality bound in top-level : and | ; omit if | Top-level kind | |
| = restriction kind in top-level : , dropping | Top-level must be exact |
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
| Field | Type | Description |
| Class identity | ||
| class_iri | string | Ontology IRI of subject class |
| class_label | string | rdfs:label of class |
| inferred | bool | True if axiom inherited via HermiT |
| axiom complexity | integer | (Section 3.3 ) |
| Correct expression | ||
| Model label (Table 3 ) | Model / HF ID | Access | Mode |
|---|---|---|---|
| GPT-5.6-sol | gpt-5.6-sol (Azure OpenAI) | Prop. | R |
| GPT-5.4 | gpt-5.4 (Azure OpenAI) | Prop. | NR |
| GPT-4o | gpt-4o (Azure OpenAI) | Prop. | NR |
| Ministral-3-14B | mistralai/Ministral-3-14B-Instruct-2512 (HF) | Open | NR |
| Qwen-3.8-27B | Qwen/Qwen3.8-27B thinking off (HF) | Open | NR |
| Qwen-3.8-27B | Qwen/Qwen3.8-27B thinking on (HF) | Open | R |