credit assignment
Determining which actions or steps in a sequence contributed to a final outcome, often for reinforcement learning.
- Papers
- 16
- Released code
- 1
- First seen
- Apr 2026
- Latest
- Sept 2026
13 papers in the last two months, against 2 in the two before.
Who is working on it
The papers
Most central to this idea first, not most recent.
- Industrycs.LG
Anchored hypergraphs provide stable credit assignment for multi-agent teams
University of Electronic Science and Technology of China, Independent Researcher · Sept 2026
- Top Universitycs.LG
Sparse verifier feedback makes broad credit assignment outperform turn targeting
Institute of Science Tokyo, Zhejiang University · Sept 2026
- Big Techcs.AIcode
Dynamic rubrics improve credit assignment for long-horizon agent training
Carnegie Mellon University, IBM Research · Sept 2026
- Independentcs.AI
Game-theoretic filtering helps agents suppress harmful memories during long tasks
Sept 2026
- Top Universitycs.AI
Privileged supervision improves action-level credit for language-model agents
Zhejiang University · Sept 2026
- Chinese Techcs.LG
Contrastive branch training improves credit assignment for tool-using language models
Alibaba Group, Harbin Institute of Technology · Aug 2026
- Top Universitycs.LG
Local milestones improve credit assignment for long-horizon language-model agents
Beijing Jiaotong University, Peking University · Aug 2026
- Independentcs.AI
Learned sepsis score tracks hourly severity without hourly outcome labels
Aug 2026
- Big Techcs.LG
Dense credit assignment from gold-answer log-probabilities improves long-horizon agent RL
University of Wisconsin–Madison, Microsoft Research · July 2026
- Chinese Techcs.CL
Separating planning from synthesis improves long-horizon search agents
Zhejiang University, Tencent · Sept 2026
- Industrycs.AI
Joint training helps small language models create and use tools
Appier AI Research, National Taiwan University · Aug 2026
- Chinese Techcs.CL
Agents learn to manage long contexts through fine-grained reinforcement learning
Tsinghua University, Tencent Youtu Lab · Sept 2026
- Top Universitycs.LG
RL with context compaction trains long-horizon agents
Tsinghua University · July 2026
- Independentcs.LG
Mixed-granularity agent graphs improve collaboration across varied tasks
Sept 2026
- AI Startupcs.CL
Chain-of-thought reasoning traces are legible but not interpretable
ETH Zurich, MIT · Sept 2026
- Big Techcs.AI
A multi-agent framework for coherent video storytelling via global optimization
Google Inc. · Apr 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.