Standard language-model reinforcement learning often gives every token in a rollout the same advantage, based on the final reward.
Chunk-level critics reduce compute for reinforcement learning on long proofs
AC2 trains a language model from critic-scored partial continuations, avoiding full rollouts when the critic is sufficiently reliable.
Independent
Kaiyue Wen · Luke Bailey · Arvind Mahankali · Tengyu Ma
Research Digest··2 min read
Wen, Bailey, Mahankali and Ma introduce Actor-Critic with Action Chunking (AC2), which assigns reinforcement-learning credit to long segments of a reasoning trajectory before a final answer is reached.
Why this paper
Independent · Released code · Part of Credit Assignment in Agentic RL, now 28 papers
In one line
AC2 exceeds GRPO's peak validation score of 18.5% using 2.5x fewer decoding FLOPs.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§