reinforcement learning with verifiable rewards
Reinforcement learning that uses rewards from objectively checkable signals like correct answers or verification feedback.
- Papers
- 8
- Released code
- 1
- First seen
- June 2026
- Latest
- Sept 2026
7 papers in the last two months, against 1 in the two before.
The papers
Most central to this idea first, not most recent.
- Top Universitycs.CLcode
Distillation Before Reinforcement Learning Improves Reasoning Model Post-Training
New York University, University of Chicago · Sept 2026
- Independentcs.AI
Five-level ladder maps reasoning models beyond direct human oversight
Sept 2026
- Chinese Techcs.LG
Contrastive branch training improves credit assignment for tool-using language models
Alibaba Group, Harbin Institute of Technology · Aug 2026
- Big Techcs.AI
Post-training can replace prompt-time schema injection for enterprise coding agents
Amazon Advertising Foundations, Amazon Web Services Agentic AI · June 2026
- Big Techcs.AI
GRPO Can Reward Lucky Guesses as If They Were Reasoning
Rochester Institute of Technology, Adobe Research · Sept 2026
- Industrycs.AI
Joint training helps small language models create and use tools
Appier AI Research, National Taiwan University · Aug 2026
- Top Universitycs.LG
Sparse verifier feedback makes broad credit assignment outperform turn targeting
Institute of Science Tokyo, Zhejiang University · Sept 2026
- Industrycs.CL
Self-evolving loop synthesizes high-quality multimodal training data
vivo AI Lab · Aug 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.