Hierarchical supervision allocation improves long-horizon model distillation
The authors propose LENS-OPD, a framework that distills a large teacher model into a smaller student agent for long-horizon tasks by hierarchically allocating supervision: first deciding how long the student should explore, then which decision to intervene on, and finally which tokens to correct. On agent benchmarks including WebShop and ScienceWorld, LENS-OPD consistently improves task success rates over vanilla on-policy distillation and strong baselines.