The author analyzes group-relative policy optimization, a reinforcement-learning approach that compares rewards among several responses to the same input without training a separate value model.
Calibrating group rewards can avoid failures from independent KL penalties
ZCPO uses each response’s relative policy drift to adjust zero-sum reward coefficients within a sampled group.
Chinese Tech
Fei Ding
Alibaba Group
Research Digest··3 min read
Ding examines why reference-policy KL regularization, which discourages a trained model from drifting too far from a baseline, can impair group-relative policy optimization.
Why this paper
From Alibaba Group
In one line
ZCPO uses conditional-KL drift to calibrate zero-sum group-relative rewards, avoiding seven failure modes that can make independent reference-policy KL regularization harmful.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§