advantage estimation
In policy optimization, a method that calculates the relative value of actions using group statistics or baselines to guide learning.
- Papers
- 4
- Released code
- 1
- First seen
- Aug 2026
- Latest
- Sept 2026
4 papers in the last two months, against 0 in the two before.
The papers
Most central to this idea first, not most recent.
- Big Techcs.AIcode
Dynamic rubrics improve credit assignment for long-horizon agent training
Carnegie Mellon University, IBM Research · Sept 2026
- Big Techcs.AI
GRPO Can Reward Lucky Guesses as If They Were Reasoning
Rochester Institute of Technology, Adobe Research · Sept 2026
- Top Universitycs.AI
Privileged supervision improves action-level credit for language-model agents
Zhejiang University · Sept 2026
- Chinese Techcs.LG
Contrastive branch training improves credit assignment for tool-using language models
Alibaba Group, Harbin Institute of Technology · Aug 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.