cs.CLOct 7, 2026

Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression

Authors: Mingqing Yuan, Xiaobo Liang, Junwei Yang, Ziwei Chen, Zeren Zhang, Hejin Wang, Yubin Wang, Juntao Li

Organizations: Soochow University · University of Cambridge · Chalmers University of Technology · Peking University · Tsinghua University · The Hong Kong University of Science and Technology

Abstract

Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction. By using the structure of rubric-guided evaluations to guide compression, LatentGRM learns compact continuous trajectories that support autonomous pairwise judgments without generating textual assessments. A separate interpreter reconstructs evaluation text from these trajectories, providing an offline view of the information retained under compression. Under matched training data and backbones, LatentGRM achieves competitive aggregate preference accuracy relative to explicit Supervised Fine-Tuning (SFT) judges at both 4B and 8B scales. Across four benchmark domains, LatentGRM-8B compresses evaluation trajectories by 8.9--9.2x and reduces total judge inference time by 6.1--7.0x at vote@5. Controlled rubric interventions show that criterion-dependent preference information is carried through the latent sequence. Together, these results demonstrate that continuous latent evaluation can substantially reduce inference cost while preserving competitive judgment quality.

Figures & tables

Appendix figures & tables19 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling

    Jun 5, 2026Xing Yue, Linjuan Wu, Daoxin Zhang +2Reward ModelingLLM Evaluation

  2. EnGRICH: Enhancing Generative Reward Modeling with Critiques from Humans

    Oct 4, 2026Xuancheng Li, Beining Wang, Haitao Li +7Reward ModelingProcess Reward Models

  3. Latent Thought Credit: Multi-Answer Credit Assignment for Latent Reasoning

    Aug 3, 2026Xuyang Zhao, Liting Zhang, Zichen Xu +4Numerical Reasoning in Language ModelsRL for Language Model Reasoning