Chunk-level critics reduce compute for reinforcement learning on long proofs

AC2 trains a language model from critic-scored partial continuations, avoiding full rollouts when the critic is sufficiently reliable.

Independent
Kaiyue Wen · Luke Bailey · Arvind Mahankali · Tengyu Ma
Research Digest··2 min read
Wen, Bailey, Mahankali and Ma introduce Actor-Critic with Action Chunking (AC2), which assigns reinforcement-learning credit to long segments of a reasoning trajectory before a final answer is reached.

Standard language-model reinforcement learning often gives every token in a rollout the same advantage, based on the final reward.

Why this paper

Independent · Released code · Part of Credit Assignment in Agentic RL, now 28 papers

In one line

AC2 exceeds GRPO's peak validation score of 18.5% using 2.5x fewer decoding FLOPs.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.