The authors studied ensemble-based self-rewarding RL, where a language model generates several responses and derives a pseudo-label or reward reference from the group, often through majority voting.
Averaging rollout contexts improves label-free language model training
GMAE estimates each response’s expected training advantage across multiple peer groups rather than trusting one randomly sampled comparison.
Chinese Tech
Yiming Wang · Yikang Liu · Qingyuan Tian · Xingyu Chen · Zhuosheng Zhang · Zhaopeng Tu · +1 more
Shanghai Jiao Tong University · Tencent
Research Digest··2 min read
Wang et al.
Why this paper
From Tencent and Shanghai Jiao Tong University
In one line
Self-rewarding LLMs become more reliable when reward estimates are marginalized over random group contexts, improving zero-label self-evolution.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§