Chen and Pang introduce Residual Advantage (RA) for reinforcement learning with verifiable rewards (RLVR).
Residual advantage uses teacher-student disagreement as bounded step reward
Chen and Pang show that centering the teacher's probability residual within each response allows step-level credit redistribution while preserving verifier outcome labels, improving reasoning performance across multiple benchmarks.
Chinese Tech
Xiaobing Chen · Zhiqi Pang
Harbin Engineering University · Tencent · Harbin Institute of Technology
Research Digest··3 min read
The authors propose Residual Advantage (RA), which treats the teacher-student probability residual as a bounded one-step reward, subtracts the student's own expected value to form an advantage, and centers it within each response.
Why this paper
From Tencent and 2 others
In one line
Residual Advantage improves RLVR by using a bounded, centered teacher-student residual as a step-level credit allocation signal.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§