The authors use a model as both student and teacher.
Selective self-distillation helps LLM judges generalize on subjective evaluations
Filtering token-level supervision by entropy shift improved out-of-distribution judgments, especially when evaluation criteria were subjective.
Big Tech
Ilgee Hong · Changlong Yu · Zhenghao Xu · Xin Liu · Yuwei Zhang · Qin Lu · +2 more
Georgia Institute of Technology · Amazon · UC San Diego
Research Digest··3 min read
Hong et al.
Why this paper
From Amazon and 2 others
In one line
Training LLM judges with self-distillation and masking high entropy shift positions improves subjective task accuracy by 2 to 9 percentage points.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§