Value-guided search improves math reasoning during RL training

APIVIS adapts Gumbel search to chunk-level reasoning, mixing direct and searched responses to create learnable rewards.

Big Tech
Shaohuai Liu · Yuning Wu · Haoran Liu · Enzo Jia · Devin Chen · Kai Wei

Amazon · Texas A&M University

Research Digest··2 min read
The authors propose APIVIS, a training-time framework that integrates finite-budget Gumbel search into reinforcement learning with verifiable rewards (RLVR) for mathematical reasoning.

The authors addressed the limitation that rollout diversity alone does not guarantee policy improvement in RLVR.

Why this paper

From Amazon and Texas A&M University

In one line

APIVIS uses value-guided search during training to improve math reasoning beyond what independent sampling achieves.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.