cs.CLOct 5, 2026

Can Language Models Learn to Reject Their Own Bad Reasoning Steps?

Authors: Siheng Xiong, Xiaoze Liu, Yiqiao Jin, Xiaoqian Wang, Jing Gao

Organizations: Georgia Institute of Technology · Purdue University

Abstract

Verifier-guided decoding can prevent harmful reasoning steps from contaminating subsequent generation, but typically relies on an external learned verifier. We ask whether a language model can instead reject its own bad reasoning steps. We define a prefix's recoverability as the probability that the frozen generator can complete it correctly. Diagnostics show that adjacent recoverability changes are often difficult to resolve with practical Monte Carlo budgets, while same-prefix candidates exhibit a sparse low-recoverability tail. We introduce Self-Step Rejection (SSR), which trains a lightweight LoRA acceptance gate on the generator backbone while keeping the base model frozen. SSR uses confidence-qualified first-passage supervision: steps before the first resolved crossing of a root-relative recoverability barrier are accepted, the crossing step is rejected, and unresolved steps and suffixes are excluded. Training combines pointwise classification, same-prefix pairwise learning, and group-relative policy refinement using final-answer correctness. At inference, SSR accepts candidates or resamples from the unchanged prefix under rejection budgets, without an external learned verifier. Across three reasoning models and five mathematical reasoning benchmarks, SSR improves macro-average accuracy over single-pass decoding by 5.4--10.1 points using 1.21--1.40x as many generated tokens, and achieves the highest macro-average accuracy among evaluated step-level methods. Full-solution scaling methods require 4.47--8.27x the single-pass token cost for comparable performance.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring

    Dec 16, 2025Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo +1Large Reasoning ModelsProgressive Reasoning

  2. From Tokens to Steps: Verification-Aware Speculative Decoding for Efficient Multi-Step Reasoning

    Apr 16, 2026Kiran Purohit, Ramasuri Narayanam, Soumyabrata PalSpeculative DecodingLLM Inference Optimization

  3. Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution

    May 19, 2026Xiaoou Liu, Tiejin Chen, Dengjia Zhang +3LLM Reasoning StrategiesReasoning Traces