The authors revisit the standard state-only value estimation in PPO for LLM reasoning tasks.
Self-privileged critic improves value estimation for RLVR
πPPO reuses verified same-prompt rollouts as contrastive evidence, boosting value quality and policy performance on math reasoning.
Chinese Tech
Kun Liang · Chenming Tang · Clive Bai · Weijie Liu · Zeyuan Liu · Qingyang Zhang · +2 more
Peking University · Tencent
Research Digest··3 min read
Liang et al.
Why this paper
From Tencent and Peking University
In one line
A self-privileged critic that uses verified same-prompt rollouts as contrastive evidence improves value estimation in RL for LLMs.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§