cs.CLSep 28, 2026

Rewarding Novel Deductions: Solver-guided Process Supervision for Logical Reasoning

Authors: Muhammad Asif Ali, Wenqing Wang, Huan Wang, Mohammad Raza

Organizations: FORTE Lab, Faculty of Science, Information Technology University, Lahore, Pakistan · College of Informatics, Huazhong Agricultural University, Wuhan, China · Qatar Computing Research Institute, Hamad Bin Khalifa University, Doha, Qatar

Abstract

Logical reasoning remains a major challenge for large language models (LLMs), particularly on structured problems that require precise constraint tracking, consistency preservation, and multi-step deduction. This challenge is especially acute for small-scale LLMs, which are more prone to producing inconsistent, redundant, or brittle reasoning trajectories. Existing approaches for improving logical reasoning largely optimize for final-answer correctness, providing only weak supervision over the intermediate reasoning process. In this work, we propose SPRING: (Solver-guided Process Rewards for Novel LogIcal ReasoNing Step Generation). SPRING uses SMT solver as a training-time verifier of intermediate reasoning steps to provide process-level supervision. It introduces the notion of a novel reasoning step, namely, a step that is logically valid, consistent with the evolving reasoning state, and not already implied by previously accepted non-contradictory deductions. Based on this solver-based assessment, it designs process rewards that encourage novel inferential progress while penalizing contradictory and uninformative reasoning steps. Evaluation across three logical reasoning benchmarks, ZebraLogic, AR-LSAT, and Knights and Knaves, and four LLMs shows that SPRING consistently outperforms base LLMs, outcome-only reward baselines, and Logic-LM. On ZebraLogic, SPRING improves puzzle accuracy by up to 49.71 and 15.43 points over the base LLM and strongest outcome-only baseline, respectively. On AR-LSAT, it improves overall accuracy by up to 64.93 and 12.14 points, respectively. On Knights and Knaves, SPRING achieves up to 93.14 puzzle accuracy and 96.05 person accuracy.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SymStep: Symbolic Step Verification for Logical Reasoning

    Jul 25, 2026Aida Usmanova, Rui Gao, Dilshod Azizov +2Constraint SatisfactionSymbol-

  2. Where Reasoning Breaks: Logic-Aware Path Selection by Controlling Logical Connectives in LLMs Reasoning Chains

    Apr 22, 2026Seunghyun Park, Yuanyuan LeiReasoning ChainReasoning Skills

  3. Discovering a Shared Logical Subspace: Steering LLM Logical Reasoning via Alignment of Natural-Language and Symbolic Views

    Apr 21, 2026Feihao Fang, My T. Thai, Yuanyuan LeiLLM Reasoning StrategiesReasoning Chain