cs.LGSep 30, 2026

Does Learning Protein Folding Generalize to Broader Reasoning?

Authors: Yong Liu, Zhanpeng Shi, Yizhou Dang, Zhongyue Zhang, Xiaoliang Shi, Zhijian Wei, Shuangjia Zheng

Organizations: Shanghai Jiao Tong University · Fudan University · Shanghai Innovation Institute · Northeastern University

Abstract

Large language models rely heavily on human text, which often conveys surface answers rather than the spatial and structural logic behind them. Protein folding is a natural testbed, because one solved structure yields thousands of exactly checkable spatial and topological statements. We ask: can learning to fold proteins teach general models reusable reasoning capabilities? To answer this, we build FoldingCorpus, a protein-derived question-answer dataset, and Fold2Reason, a recipe that post-trains on it through two complementary signals: discrete structural answers predicted via the model's native language head, and continuous 3D geometry decoded from the same shared representations. On FoldBench, Fold2Reason achieves structure prediction scores 2.7 to 3.5 times those of Qwen3.5-9B. Beyond protein structure prediction, it improves performance on all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33% (+3.23 pp), with positive gains on all 10 benchmarks, while matched controls built from random, synthetic, and shuffled structure yield substantially smaller or negative gains. Our work shows that non-linguistic, structure-dense scientific data can systematically improve broad reasoning in language models, making a solved scientific problem a practical source of post-training supervision.

Figures & tables

Appendix figures & tables31 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

    Jul 8, 2026Chen Tang, Yizhou Wang, Jianyu Wu +26Scientific DiscoveryReasoning Skills

  2. How Post-Training Shapes Biological Reasoning Models

    Jun 15, 2026Lukas Fesser, Hanlin Zhang, Michelle M. Li +5Large Reasoning ModelsPost-Training

  3. Wrong Prediction, Right Answer: Recovering Evidence from Collapsed LLM Sequence Scores

    Aug 31, 2026Qiyao Yan, Chenpeng Wang, Liangming PanLLM Reasoning StrategiesReasoning Benchmark