group relative policy optimization
A reinforcement learning method that optimizes policy by computing advantages based on group-level reward statistics.
- Papers
- 8
- Released code
- 3
- First seen
- June 2026
- Latest
- Sept 2026
7 papers in the last two months, against 1 in the two before.
Who is working on it
The papers
Most central to this idea first, not most recent.
- Chinese Techcs.CV
Video-grounded prompt planning improves long-form text-to-video generation quality
Nanjing University, Wan Team, Alibaba Group · Sept 2026
- Big Techcs.AI
GRPO Can Reward Lucky Guesses as If They Were Reasoning
Rochester Institute of Technology, Adobe Research · Sept 2026
- Big Techcs.AIcode
Dynamic rubrics improve credit assignment for long-horizon agent training
Carnegie Mellon University, IBM Research · Sept 2026
- Top Universitycs.AI
Privileged supervision improves action-level credit for language-model agents
Zhejiang University · Sept 2026
- Top Universitycs.AI
Step-level checks curb unsafe agent actions with little utility loss
Shanghai Artificial Intelligence Laboratory, Beihang University · Aug 2026
- Chinese Techcs.CV
Conversational image editing agent learns interpretable tool use
Shanghai Innovation Institution, Huawei Technologies Ltd. · June 2026
- Big Techcs.CLcode
Paper-derived rubrics improve AI generation of scientific research plans
Zhejiang University, Apple · Sept 2026
- Top Universitycs.AIcode
Reinforcement learning strengthens policy invocation for agent safety judgments
Zhejiang University, Zhongguancun Academy · Aug 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.