Averaging rollout contexts improves label-free language model training

GMAE estimates each response’s expected training advantage across multiple peer groups rather than trusting one randomly sampled comparison.

Chinese Tech
Yiming Wang · Yikang Liu · Qingyuan Tian · Xingyu Chen · Zhuosheng Zhang · Zhaopeng Tu · +1 more

Shanghai Jiao Tong University · Tencent

Research Digest··2 min read
Wang et al.

The authors studied ensemble-based self-rewarding RL, where a language model generates several responses and derives a pseudo-label or reward reference from the group, often through majority voting.

Why this paper

From Tencent and Shanghai Jiao Tong University

In one line

Self-rewarding LLMs become more reliable when reward estimates are marginalized over random group contexts, improving zero-label self-evolution.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.