Token-level rewards make critic-free reasoning training faster and more accurate

BoT-GRPO assigns localized advantages while balancing contributions from responses of different lengths, improving convergence on code generation and mathematics.

Big Tech
Yingxiang Yang · Weihang Xiao · Zhunxuan Wang · Joshua Flashner · Niresh Agarwal

Amazon AGI · Virginia Tech

Research Digest··2 min read
Yang and colleagues introduce Bag-of-Tokens Group Relative Policy Optimization, a modification of GRPO that uses token-level process rewards without learning a value network.

Standard GRPO gives every token in a sampled response the same sequence-level advantage, even when only particular steps helped or harmed the result.

Why this paper

From Amazon AGI and Virginia Tech

In one line

BoT-GRPO uses length-normalized token rewards to train reasoning models faster than GRPO, while improving React compilation, visual quality, and AIME results without a critic.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (4 noted)
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.