Judging in Latent Space: Efficient Generative Reward Modeling via Semantics-Preserving Compression
Organizations: Soochow University · University of Cambridge · Chalmers University of Technology · Peking University · Tsinghua University · The Hong Kong University of Science and Technology
Abstract
Reward modeling often requires jointly representing and reasoning over multiple evaluation criteria, yet verbalizing this process token by token can incur substantial inference cost. Recent work on latent reasoning suggests that continuous states may support this computation more compactly. We introduce LatentGRM, a latent evaluation framework built on semantic chunking, compression, and reconstruction. By using the structure of rubric-guided evaluations to guide compression, LatentGRM learns compact continuous trajectories that support autonomous pairwise judgments without generating textual assessments. A separate interpreter reconstructs evaluation text from these trajectories, providing an offline view of the information retained under compression. Under matched training data and backbones, LatentGRM achieves competitive aggregate preference accuracy relative to explicit Supervised Fine-Tuning (SFT) judges at both 4B and 8B scales. Across four benchmark domains, LatentGRM-8B compresses evaluation trajectories by 8.9--9.2x and reduces total judge inference time by 6.1--7.0x at vote@5. Controlled rubric interventions show that criterion-dependent preference information is carried through the latent sequence. Together, these results demonstrate that continuous latent evaluation can substantially reduce inference cost while preserving competitive judgment quality.
Figures & tables
| Model | RewardBench | IF Evaluation | RM-Bench | RewardBench2 | HS3 | Avg.-8 | |||
|---|---|---|---|---|---|---|---|---|---|
| Chat | Hard | PPE | IFBench | Chat | Prec. IF | Focus | |||
| White-box judge/reward LLMs (for reference only) | |||||||||
| JudgeLRM-7B | 93.2 | 56.4 | 47.9 | 47.5 | 55.0 | 29.7 | 39.9 | 60.5 | 53.7 |
| RRM-7B | 88.0 | 70.6 | 52.5 | 55.2 | 59.8 | 29.4 | 63.9 | 63.7 | 60.4 |
| RM-R1-7B (Qwen-2.5-Inst) | 94.4 | 72.3 | 56.1 | 57.8 | 65.8 | 35.6 | 80.0 | 69.3 | 66.4 |
| RM-R1-7B (DeepSeek-Dist) | 86.2 | 67.0 | 52.5 | 56.8 | 62.1 | 23.8 | 56.9 | 63.3 | 58.6 |
| RewardBench | RewardBench2 | |||
| Method | Chat | Hard | Prec. IF | Focus |
| vote@1 | ||||
| Fixed-Rate Chunking | 90.5 | 69.0 | 37.2 | 83.3 |
| Semantic Chunking | 88.7 | 70.5 | 37.9 | 84.6 |
| vote@5 | ||||
| Fixed-Rate Chunking | 90.6 | 70.8 | 39.5 | 85.8 |
| Domain ( ) | Model | Steps | Compression | HF time (s) | HF speedup | vLLM time (s) | vLLM speedup |
|---|---|---|---|---|---|---|---|
| RB Chat (716) | Rubric-RM-8B | 659.60 | 8196 | 539.70 | |||
| LatentGRM-8B | 74.38 | 1002 | 87.76 | ||||
| RB Chat Hard (912) | Rubric-RM-8B | 582.09 | 11088 | 589.58 | |||
| LatentGRM-8B | 64.42 | 1308 | 96.44 | ||||
| RB2 Precise IF (960) | Rubric-RM-8B | 720.89 | 13416 | 812.93 | |||
| LatentGRM-8B | 78.59 | 1524 | 126.99 |
| Readout input | State Macro-F1 | Vector Exact | Rubric Semantic | ROUGE-L |
|---|---|---|---|---|
| Prompt-only | 40.20 | 18.61 | 56.42 | 63.36 |
| Replay-hidden | 69.42 | 33.03 | 61.29 | 68.22 |
| Latent Trace Interpreter | 98.93 | 93.98 | 74.65 | 79.24 |
| Feedback interface | Repair | Regression | Full | Criterion sat. |
|---|---|---|---|---|
| Prompt-only | 54.07 | 5.20 | 34.17 | 76.28 |
| Rubric-RM trajectory | 78.71 | 4.65 | 56.67 | 87.78 |
| Latent only (soft top-10) | 63.32 | 5.11 | 41.67 | 80.54 |
| Interpreter-restored trajectory | 80.49 | 1.86 | 57.50 | 90.12 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Boundary type | Base cost |
|---|---|
| Observed sentence end | |
| Complete criterion-header end or predicate-bearing clause edge | |
| Soft punctuation, such as a comma, colon, or semicolon | |
| Lexical separator within code or mathematics | |
| Ordinary word boundary | |
| Internal position of a criterion header |
| Explicit tokens per latent | Chunk count | Share (%) |
|---|---|---|
| 4 | 37,514 | 1.08 |
| 5 | 223,876 | 6.47 |
| 6 | 399,275 | 11.55 |
| 7 | 572,989 | 16.57 |
| 8 | 1,053,395 | 30.46 |
| 9 | 523,161 | 15.13 |
| Setting | Value |
|---|---|
| Initialization | LatentGRM-8B Stage 1 decoder; frozen backbone |
| Training data | 28,430 examples; 512 validation examples; seed 42 |
| Trajectory view | Full trajectory in forward order; one view per example |
| LTI representation | Probability-weighted top- input embeddings |
| LoRA | Rank 32; ; dropout 0.05; all attention and MLP projections |
| Training duration | 2 epochs |
| Component | Text from the accepted example |
|---|---|
| Request | Who created the Superman cartoon character? |
| Response A | Superman, the iconic comic book superhero, was created by writer Jerry Siegel and artist Joe Shuster. Superman first appeared in Action Comics #1, which was published by Detective Comics, Inc. (later DC Comics) in June 1938. The character’s immense popularity established him as one of the most enduring and recognizable figures in the superhero genre. |
| Response B | The Superman cartoon character was created by Jerry Siegel and Joe Shuster in 1938. |
| Original | 3. The response should be concise and avoid unnecessary elaboration. [Principle] |
| Paraphrase | 3. The response should be brief and avoid including extraneous information. [Principle] |
| Counterfactual | 3. The response should provide thorough and enriching details that go beyond the minimum facts required to answer. [Principle] |
| Component | Text from the accepted example |
|---|---|
| Request | Generate an advertisement slogan that promotes healthy eating. |
| Response A | “Get Fit, Fuel Up - Find Balance with Our App!” |
| Response B | Eat healthy, feel healthy! |
| Original | 5. The response should employ vivid, engaging imagery or metaphors to capture attention. [Principle] |
| Paraphrase | 5. The response should use striking, imaginative language or figurative expressions to draw attention. [Principle] |
| Counterfactual | 5. The response should use simple, literal language that directly communicates its message without figurative embellishment. [Principle] |
| RewardBench | RewardBench2 | ||||
|---|---|---|---|---|---|
| Model | Votes | Chat | Chat Hard | Precise IF | Focus |
| DirectJudge | 1 | 90.92 | 67.21 | 35.94 | 85.15 |
| 5 | 91.62 | 67.43 | 41.56 | 86.57 | |
| 9 | 90.78 | 67.56 | 41.25 | 86.47 | |
| Rubric-RM | 1 | 89.80 | 69.80 | 31.90 | 77.90 |
| 5 | 91.80 | 70.50 | 45.00 | 81.10 | |
| Model | Mean prompt positions | Output positions | Total inference time (s) | Throughput (positions/s) |
|---|---|---|---|---|
| Rubric-RM-8B | 913.8 | 81,400 | 46.469 | 1,751.7 |
| LatentGRM-8B | 914.8 | 81,400 | 45.714 | 1,780.6 |
| Readout input | Met F1 | Not-Met F1 | Partial F1 |
|---|---|---|---|
| Prompt-only | 59.95 | 50.21 | 10.44 |
| Replay-hidden | 94.45 | 79.46 | 34.36 |
| Latent Trace Interpreter | 99.64 | 98.68 | 98.46 |
| Setting | Value |
|---|---|
| Backbones / training records | Qwen3-4B and Qwen3-8B; 35,612 records |
| Hardware per run | NVIDIA RTX 6000D (85,651 MiB per GPU) |
| Precision / distributed training | bfloat16; DeepSpeed ZeRO Stage 2; gradient checkpointing |
| Maximum sequence length | 6,144 tokens |
| Compression rate / latent support | ; top- ; projection temperature |
| Semantic Chunking constraints | hard 4–16 tokens; preferred 6–10 tokens; local search radius 4 tokens (5 for lexical rescue) |