Credit assignment that prunes redundant steps improves reasoning accuracy and efficiency

The authors' RECAP method reshapes GRPO advantages into step-specific updates using semantic dependency graphs and answer-directed progress, shortening reasoning traces without hurting accuracy.

Big Tech
Yuqing Zhou · Hong Wang · Manqing Mao · Zhuoer Wang · Samson Koelle · Jie Yuan · +5 more

George Mason University · Amazon, Inc

Research Digest··2 min read
The authors introduce RECAP, a redundancy-aware credit assignment method that assigns step-level rewards based on both a step's downstream role in the reasoning structure and its progress toward the correct answer.

The authors propose RECAP (REdundancy-aware Credit Assignment via Propagation), which combines two signals: structural responsibility, measured by propagating credit backward through an LLM-annotated semantic dependency graph from the final answer node, and step efficacy, measured by the change in gold-answer log-likelihood as each step is added.

Why this paper

From Amazon, Inc and George Mason University · Part of Credit Assignment in Agentic RL, now 30 papers

In one line

RECAP assigns credit based on downstream dependency and correct-answer progress, improving accuracy and reducing reasoning tokens in large models.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.