cs.LGSep 30, 2026

MatrixReward: Reward from Rubric Matrix for Open-Ended Generation

Authors: Zihan Shen, Qi Liu, Zixuan Yang, Yiqun Chen, Chenglong Zhao, Xiaozhao Wang, Lei He

Organizations: Zhejiang University · Qwen Business Unit of Alibaba · Renmin University of China

Abstract

Open-ended query generation lacks standard answers, thus necessitating an effective reward mechanism. Pointwise scoring rubrics provide limited information about the relative quality of sample answers under the same prompt; merging multiple rubric judgments into a single score may also mask the differences between these answers. We propose MatrixReward, which constructs rewards from a rollout-by-rubric win-rate matrix obtained by comparing every pair of sampled responses under each rubric. The spread of each matrix column captures how strongly that rubric distinguishes the current rollouts, while correlations between columns reveal rubric repetition; together, these statistics yield data-dependent rubric weights. We combine these weights with the prior weights of rubrics. After column normalization and weighting, the observed per-rubric maxima and minima define positive and negative ideal profiles. Each rollout's distances to these two ideals determine its relative-closeness quality reward. Evaluated using Qwen3-8B on four open-ended query-answering benchmarks, MatrixReward achieves an average score of 63.02, outperforming the strongest baseline by approximately 2.0%. These results support the idea that matrices derived from relative comparisons can be used to construct rewards more reasonably for open-ended generative reinforcement learning.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation

    Sep 29, 2026Zixuan Yang, Yiqun Chen, Qi Liu +6Group-Based Reinforcement Learning

  2. Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling

    Jul 2, 2026Dazhi Fu, Jiuding Yang, Yiwen Guo +1Task-Specific RubricsProgress Reward Modeling

  3. Prompt-Level Reward Specifications for Open-Ended Post-Training

    May 28, 2026Zijun Weng, Xiaohui Hu, Shuangyong Song +3Post-TrainingOffline Reinforcement Learning