Divergence choice in distillation controls student model entropy

Forward KL inflates entropy while reverse KL deflates it, acting as an implicit regularizer.

Top University
Nicolas Zucchet · Scott W. Linderman

Stanford University

Research Digest··2 min read
The authors decouple the sampling distribution from the divergence used in language model distillation.

Zucchet and Linderman study distillation from an entropic perspective, disentangling the choice of data generation (on- vs.

Why this paper

From Stanford University

In one line

Divergence choice in distillation controls student entropy, with forward KL inflating and reverse KL deflating it.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.