For each prompt, the authors sampled multiple responses, or rollouts, then compared every pair under each evaluation rubric.
Relative rubric comparisons improve rewards for open-ended language model training
MatrixReward turns pairwise, rubric-specific judgments into adaptive rewards that distinguish sampled responses without requiring reference answers.
Chinese Tech
Zihan Shen · Qi Liu · Zixuan Yang · Yiqun Chen · Chenglong Zhao · Xiaozhao Wang · +1 more
Zhejiang University · Renmin University of China · Qwen Business Unit of Alibaba
Research Digest··2 min read
Shen and colleagues propose MatrixReward, a reward construction method for reinforcement learning on open-ended tasks where no single correct answer exists.
Why this paper
From Qwen Business Unit of Alibaba and 2 others
In one line
MatrixReward uses pairwise rubric comparisons to build a win-rate matrix and derive distance-based rewards for open-ended RL.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 8B)
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§