Standard GRPO gives every token in a sampled response the same sequence-level advantage, even when only particular steps helped or harmed the result.
Token-level rewards make critic-free reasoning training faster and more accurate
BoT-GRPO assigns localized advantages while balancing contributions from responses of different lengths, improving convergence on code generation and mathematics.
Big Tech
Yingxiang Yang · Weihang Xiao · Zhunxuan Wang · Joshua Flashner · Niresh Agarwal
Amazon AGI · Virginia Tech
Research Digest··2 min read
Yang and colleagues introduce Bag-of-Tokens Group Relative Policy Optimization, a modification of GRPO that uses token-level process rewards without learning a value network.
Why this paper
From Amazon AGI and Virginia Tech
In one line
BoT-GRPO uses length-normalized token rewards to train reasoning models faster than GRPO, while improving React compilation, visual quality, and AIME results without a critic.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (4 noted)
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§