cs.CLMar 18, 2026

RAWR: Reward Assignment Without Rollouts in Verifiable Domains

Authors: Corentin Royer, Anna Hedström, Debarun Bhattacharjya, Gaetano Rossiello, Andrea Giovannini, Mennatallah El-Assady

Organizations: International Business Machines · Department of Computer Science, ETH Zurich, 8092 Zurich, Switzerland · ETH AI Center · Lirio · Department of Computer Science, ETH Zurich

Abstract

Understanding and evaluating multi-step reasoning in LLMs at the level of individual steps remains a key challenge. Process reward models (PRMs) provide a solution by scoring each step, enabling fine-grained supervision and improved reliability. However, training them requires costly human annotation or computationally intensive rollout-based labeling. To solve this, we introduce MCNIG, a scalable method for automatically labeling the quality of individual reasoning steps in any verifiable domain. Its step score, net information gain (NetIG), improves upon single-reference information gain (IG) by comparing the most-supported correct answer against the most-supported incorrect one, yielding a robust signal even for long and structured outputs like code and SQL, where IG fails. We show that the signal produced by MCNIG correlates with human judgments of step quality, and we apply MCNIG labels to train PRMs that achieve the best average best-of-K accuracy across eight benchmarks spanning mathematics, code generation, text-to-SQL, and scientific QA. Crucially, MCNIG generates no rollouts, cutting labeling complexity to O(N) and making it up to X times cheaper than rollout-based methods at comparable label quality, which makes large-scale process supervision practical.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Step-wise Rubric Rewards for LLM Reasoning

    May 17, 2026Weichu Xie, Haozhe Zhao, Wenpu Liu +15Reinforcement Learning With Verifiable RewardRubric-Based Scoring

  2. ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth

    Dec 2, 2025Salman Rahman, Sruthi Gorantla, Arpit Gupta +3Process Reward ModelMathematical Reasoning Benchmarks

  3. Rethinking Reward Models for Multi-Domain Test-Time Scaling

    Oct 1, 2025Dong Bok Lee, Seanie Lee, Sangwoo Park +12Process Reward ModelLLM Reasoning Strategies