On-policy distillation (OPD) trains a student model on its own generated trajectories with dense teacher supervision, but long rollouts are expensive and teacher guidance degrades as student prefixes drift into regions of low compatibility.
Adaptive per-trajectory stopping speeds on-policy distillation by up to 7.5x
Flash-OPD replaces shared rollout horizons with exact cumulative compatibility checks, preserving accuracy while cutting rollout cost.
Chinese Tech
Wei Chen · Junle Chen · Yitong Yang · Zhaoyang Xu · Jiaxin Lin · Yuxuan Liang · +3 more
Tencent Hy · HKUST(GZ) · HKUST
Research Digest··3 min read
The authors propose Flash-OPD, a method for on-policy distillation that stops each training trajectory individually based on observed teacher-student compatibility events rather than using a fixed rollout length.
Why this paper
From Tencent Hy and 2 others
In one line
Flash-OPD speeds up on-policy distillation by 2.2 to 7.5 times while maintaining or improving accuracy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§