Credit Assignment in Agentic RL
Methods for distributing credit over long trajectories in language-model agent reinforcement learning when only sparse terminal rewards are available.
17 papers · 3 months
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.
Results across this thread
8 reported results from the papers in this thread.
The table is for subscribers. Subscribe to see every reported number side by side.
How this thread developed
June 2026 · IBM Research
RL with live server state boosts multi-step tool use in LLMs
Presents an RL framework for multi-step tool use that addresses credit assignment with live server state.
4 further papers
Aug 2026 · Fudan University, Tencent
Search agents improve longer when their critics evolve alongside them
Introduces CAFE, which alternates search and critic feedback, enabling credit assignment via online and offline RL.
3 further papers
Sept 2026 · Carnegie Mellon University, IBM Research
Dynamic rubrics improve credit assignment for long-horizon agent training
DRACO uses dynamic rubrics and stepwise rewards to improve credit assignment for long-horizon agent training when success is not programmatically checkable.
released code
4 further papers
Sept 2026 · University of Maryland, AWS AI Labs
Predicting environment observations during fine-tuning improves later agent exploration
ActObs predicts environment observations during fine-tuning, improving later exploration—a form of auxiliary supervision that supports credit assignment over agent trajectories.
Sept 2026 · Xi’an Jiaotong University, National Key Laboratory for Multimedia Information Processing
Handheld demonstrations improve robot policies without repeated robot execution
Uses handheld demonstrations to refine policies, addressing credit assignment in manipulation tasks without repeated robot execution.
Sept 2026 · UT Austin, Autel US
Targeted subtask reinforcement learning improves long-horizon robot manipulation
Targets subtask-level credit assignment in long-horizon RL, learning corrections for specific failure-prone subtasks while freezing the base policy.
6 of 17 papers shown