Adaptive per-trajectory stopping speeds on-policy distillation by up to 7.5x

Flash-OPD replaces shared rollout horizons with exact cumulative compatibility checks, preserving accuracy while cutting rollout cost.

Chinese Tech
Wei Chen · Junle Chen · Yitong Yang · Zhaoyang Xu · Jiaxin Lin · Yuxuan Liang · +3 more

Tencent Hy · HKUST(GZ) · HKUST

Research Digest··3 min read
The authors propose Flash-OPD, a method for on-policy distillation that stops each training trajectory individually based on observed teacher-student compatibility events rather than using a fixed rollout length.

On-policy distillation (OPD) trains a student model on its own generated trajectories with dense teacher supervision, but long rollouts are expensive and teacher guidance degrades as student prefixes drift into regions of low compatibility.

Why this paper

From Tencent Hy and 2 others

In one line

Flash-OPD speeds up on-policy distillation by 2.2 to 7.5 times while maintaining or improving accuracy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.