on-policy distillation
- Papers
- 17
- Released code
- 2
- First seen
- Sept 2026
- Latest
- Oct 2026
17 papers in the last two months, against 0 in the two before.
Who is working on it
The papers
Most central to this idea first, not most recent.
- Big Techcs.AI
Training agents around decisive errors improves prevention and recovery
Princeton University, NVIDIA · Oct 2026
- Chinese Techcs.CL
Teacher-guided training improves specialization and coordination in multi-agent models
University of Science and Technology of China, Tencent · Sept 2026
- Chinese Techcs.CL
Hierarchical supervision allocation improves long-horizon model distillation
Ant Group, Alibaba International Digital Commerce Group · Sept 2026
- Chinese Techcs.LGcode
Token-level teacher routing beats prompt-level routing in multi-teacher distillation
Shanghai Jiao Tong University, GAIR · Sept 2026
- Chinese Techcs.LG
Agent distillation improves when supervision follows outcomes, not model disagreement
Qwen Large Model Application Team, Alibaba, Peking University · Sept 2026
- Chinese Techcs.LG
Execution graph helps off-the-shelf teachers supervise student agents better
Yuanbao Team, Tencent, Tsinghua University · Sept 2026
- Chinese Techcs.LG
Online distillation cuts needless verbosity in RL-trained models
ByteDance Seed, Tongji University · Oct 2026
- Chinese Techcs.LG
Ablating individual constraints improves distillation for complex instruction following
Alibaba Group · Sept 2026
- Big Techcs.LG
Attention distillation complements token-level supervision for reasoning models
ACI PLC, University of Dhaka · Sept 2026
- Big Tech
Co-evolving teacher and student models improves mathematical reasoning
Meta AI, University of California, Riverside · Sept 2026
- Chinese Techcs.CV
Real-time GUI feedback improves training for computer-use agents
Zhejiang University, Ant Group · Oct 2026
- Chinese Techcs.AI
Self-distillation trains GUI agents for longer, memory-dependent tasks
Institute of Information Engineering, Chinese Academy of Sciences, Tencent · Sept 2026
- Chinese Techcs.CV
Real-time GUI feedback improves training for computer-use agents
Zhejiang University, Ant Group · Oct 2026
- Chinese Techcs.LG
Training agents on their own trajectories accelerates reward-based learning
Zhejiang University, Ant Healthcare · Oct 2026
- Top Universitycs.CLcode
Distillation Before Reinforcement Learning Improves Reasoning Model Post-Training
New York University, University of Chicago · Sept 2026
- Chinese Techcs.CL
Terminal agents verify often but miss and mishandle many errors
Northeastern University, Inclusion AI, Ant Group · Oct 2026
- Chinese Techcs.AI
PlanGuard detects physical risks across complete embodied-agent plans
Anhui Province Key Laboratory of Digital Security, The University of Hong Kong · Sept 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.