cs.AISep 29, 2026

Learning to Prove, Not Just to Answer: Reinforcement Learning from Formal Verification for Natural-Language Logical Reasoning

Authors: Qili Zhang, Qianren Mao, Hanze Cai, Kaiming Zhao, Yuening He, Xihan Lei, Yashuo Luo, Hanwen Hao, +6 more

Organizations: Zhongguancun Laboratory · Beihang University · Nanyang Technological University · Hong Kong Polytechnic University

Abstract

Large language models (LLMs) are increasingly deployed for natural-language logical reasoning, where the final answer is easy to check but the proof behind it is not. In natural-language logical reasoning, an intermediate conclusion should follow from its premises, and the resulting derivation should support the final answer. Existing methods lack machine-checkable verification of intermediate conclusions and answer-supporting proof dependencies, so they may assign credit to invalid or answer-irrelevant steps. We propose Proof-R1, an RL framework from formal verification that trains LLMs to construct verifiable proofs for natural-language logical reasoning. Proof-R1 admits a generated conclusion into the verified proof state only when the corresponding reasoning action satisfies the proof obligations through UNSAT-based machine-checkable formal verification. Proof-R1 also recovers the answer-supporting dependency closure to trace the proof structure of the final answer and align outcome credit with the proof dependencies. Experiments demonstrate that Proof-R1 improves answer accuracy across three logical reasoning benchmarks and four backbone models and outperforms training-free agents and training-based methods in terms of reasoning-process verifiability.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Logic-Regularized Verifier Elicits Reasoning from LLMs

    May 7, 2026Xinyu Wang, Changzhi Sun, Lian Cheng +4LLM Reasoning StrategiesReasoning Paths

  2. ProofVerifier: A Scalable, Diversity-Driven Framework for Natural-Language Proof Verification

    Feb 2, 2026Haotong Yang, Zitong Wang, Shijia Kang +7ProofLarge Language Model Reliability

  3. Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models

    Jul 16, 2026Jungseob Lee, Seungyoon Lee, Suhyune Son +4LLM Reasoning StrategiesReasoning Skills