Selective self-distillation helps LLM judges generalize on subjective evaluations

Filtering token-level supervision by entropy shift improved out-of-distribution judgments, especially when evaluation criteria were subjective.

Big Tech
Ilgee Hong · Changlong Yu · Zhenghao Xu · Xin Liu · Yuwei Zhang · Qin Lu · +2 more

Georgia Institute of Technology · Amazon · UC San Diego

Research Digest··3 min read
Hong et al.

The authors use a model as both student and teacher.

Why this paper

From Amazon and 2 others

In one line

Training LLM judges with self-distillation and masking high entropy shift positions improves subjective task accuracy by 2 to 9 percentage points.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.