cs.AISep 29, 2026

CARAT: Do Materials LLMs Reason or Recite?

Authors: Jiajun Wu, Jian Yang, Zixiang Ni, Zhenzhu Li, Bin Chong

Organizations: Beihang University · Xi’an Jiaotong University · Imperial College London · Peking University

Abstract

When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names each structural relation separately in GraphSpace, and adds matched fine-tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims. First, on the benchmark's hardest families the grounded view is worth 17.3 points over formula inputs. Second, we turn that scrutiny on ourselves. GraphSpace beats a plain periodic graph by 19.3 points, but that margin is two effects at once: where the plain rendering carries everything the question needs it is 1.96 points, and where it omits those fields entirely, 46.7 points. The headline mostly measures what the baseline lacked, not how evidence is presented. Third, we attack our own benchmark. A rule that skips the link and reads the list directly answers four of seven hardened families, so we rebuilt it until eleven such shortcuts sat near chance. The frozen model quotes that link yet answers the same when we redirect it, on 95.6% of paired cases: it repeats the relation without using it. After matched supervision it reaches 99.8%, and deleting the link drops it to 23.4%, below the 27.0% the best shortcut reaches: both steps are learnable.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

    May 28, 2026Wanhao Liu, Jiaqing Xie, Qian Tan +10Materials ScienceMaterials

  2. Visualizing Graph-to-Answer Mechanism Recovery in Materials-Science Hypothesis Generation

    Aug 4, 2026Shashwat Sourav, Subhadeep Pal, Markus J. Buehler +4HypothesisSynthesis

  3. Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis

    Apr 15, 2026Zipeng Ling, Shuliang Liu, Seonil Son +4Reasoning TracesChain-of-Thought Reasoning