Sun et al.
Hierarchical supervision allocation improves long-horizon model distillation
The LENS-OPD method outperforms dense token-level supervision by dynamically selecting which student behaviors to correct based on future impact and learnability.
Chinese Tech
Yuhao Sun · Binrui Wu · Zhuoer Xu · Ming Wen · Haoxiang Xu · Bin Chen · +2 more
Ant Group · Alibaba International Digital Commerce Group · University of Science and Technology of China · Peking University · University of Electronic Science and Technology of China
Research Digest··3 min read
The authors propose LENS-OPD, a framework that distills a large teacher model into a smaller student agent for long-horizon tasks by hierarchically allocating supervision: first deciding how long the student should explore, then which decision to intervene on, and finally which tokens to correct.
Why this paper
From Ant Group and 4 others · Part of Credit Assignment in Agentic RL, now 28 papers
In one line
Hierarchical supervision allocation via Locate-Validate-Refine improves long-horizon on-policy distillation for LLM agents.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§