Zucchet and Linderman study distillation from an entropic perspective, disentangling the choice of data generation (on- vs.
Divergence choice in distillation controls student model entropy
Forward KL inflates entropy while reverse KL deflates it, acting as an implicit regularizer.
Top University
Nicolas Zucchet · Scott W. Linderman
Stanford University
Research Digest··2 min read
The authors decouple the sampling distribution from the divergence used in language model distillation.
Why this paper
From Stanford University
In one line
Divergence choice in distillation controls student entropy, with forward KL inflating and reverse KL deflating it.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§