The authors addressed the limitation that rollout diversity alone does not guarantee policy improvement in RLVR.
Value-guided search improves math reasoning during RL training
APIVIS adapts Gumbel search to chunk-level reasoning, mixing direct and searched responses to create learnable rewards.
Big Tech
Shaohuai Liu · Yuning Wu · Haoran Liu · Enzo Jia · Devin Chen · Kai Wei
Amazon · Texas A&M University
Research Digest··2 min read
The authors propose APIVIS, a training-time framework that integrates finite-budget Gumbel search into reinforcement learning with verifiable rewards (RLVR) for mathematical reasoning.
Why this paper
From Amazon and Texas A&M University
In one line
APIVIS uses value-guided search during training to improve math reasoning beyond what independent sampling achieves.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§