agentic reinforcement learning
Training approach that lets LLM-based agents learn multi-step tool use and planning via reinforcement learning, often with dense or per-tool rewards.
- Papers
- 26
- Released code
- 4
- First seen
- Mar 2026
- Latest
- Sept 2026
18 papers in the last two months, against 7 in the two before.
Who is working on it
The papers
Most central to this idea first, not most recent.
- Independentcs.AI
Closed-loop development improves mobile agents across planning and tool use
Sept 2026
- Industrycs.AI
Co-trained agent roles improve tool-based reasoning and verification
Waseda University, Adelaide University · Sept 2026
- Chinese Techcs.CLcode
Self-retiring distillation improves reinforcement learning for multi-turn agents
Zhejiang University, Alibaba Group · Sept 2026
- Big Techcs.CV
A single visual agent spans mobile, desktop, web, and tool use
Apple · Sept 2026
- Top Universitycs.CV
Reinforcement learning trains video AI agents to use external tools effectively
Princeton University, Stanford University · Sept 2026
- Industrycs.AI
Joint training helps small language models create and use tools
Appier AI Research, National Taiwan University · Aug 2026
- Chinese Techcs.LG
Shared-prefix training accelerates reinforcement learning for hybrid-attention agents
Ant Group · Sept 2026
- Big Techcs.LG
Dense credit assignment from gold-answer log-probabilities improves long-horizon agent RL
University of Wisconsin–Madison, Microsoft Research · July 2026
- Big Techcs.LG
A lightweight PyTorch-native framework matches Megatron-based agentic RL training performance.
NVIDIA · July 2026
- Big Techcs.LG
Predicting environment observations during fine-tuning improves later agent exploration
University of Maryland, AWS AI Labs · Sept 2026
- Industrycs.AI
Synthetic rewards train agents to diagnose simulated advertising anomalies
Independent Researchers · Sept 2026
- Top Universitycs.AI
Privileged supervision improves action-level credit for language-model agents
Zhejiang University · Sept 2026
- Chinese Techcs.LG
Contrastive branch training improves credit assignment for tool-using language models
Alibaba Group, Harbin Institute of Technology · Aug 2026
- Top Universitycs.LG
RL with context compaction trains long-horizon agents
Tsinghua University · July 2026
- Chinese Techcs.AIcode
Evolutionary training harness co-evolves with LLM policies for RL
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Tongyi Lab , Alibaba Group · June 2026
- Big Techcs.AIcode
Dynamic rubrics improve credit assignment for long-horizon agent training
Carnegie Mellon University, IBM Research · Sept 2026
- Industrycs.CR
Graph-based policy constrains LLM agents for topology-aware incident response
Indiana University · Sept 2026
- Big Techcs.AI
New framework lets you train AI agents inside the same harness systems they use at inference
Columbia University, Dartmouth College · July 2026
- Chinese Techcs.CL
Separating planning from synthesis improves long-horizon search agents
Zhejiang University, Tencent · Sept 2026
- Chinese Techcs.CL
Agents learn to manage long contexts through fine-grained reinforcement learning
Tsinghua University, Tencent Youtu Lab · Sept 2026
- Independentcs.CLcode
A 2.8-trillion-parameter open MoE model approaches frontier performance.
July 2026
- Chinese Techcs.AI
Three-stage training instills world model planning into LLM agents
Fudan University, Shanghai Innovation Institute · June 2026
- Independentcs.LG
Mixed-granularity agent graphs improve collaboration across varied tasks
Sept 2026
- Chinese Techcs.AI
Evolving terminal environments keeps training tasks challenging as agents improve
Hunyuan Team, Tencent · Sept 2026
- Chinese Techcs.CV
Unified visual-generation agentic model outperforms larger closed-source models
Tencent Hunyuan, Hong Kong University of Science and Technology · Mar 2026
- Top Universitycs.AI
Enumerating Tool Choices Beats Sampling Them in Genomic Reasoning
The Chinese University of Hong Kong, Shenzhen, Shenzhen Future Network of Intelligence Institute (FNii-Shenzhen) · Sept 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.