Online distillation cuts needless verbosity in RL-trained models

Length Self-Distillation reduces excess response length on easy queries without sacrificing accuracy on hard ones.

Chinese Tech
Xu Wan · Wenyue Xu · Shengjie Zhao · Mingyang Sun

ByteDance Seed · Tongji University · Peking University

Research Digest··3 min read
The authors identify and quantify the 'length-scaling tax'—excess response length on already-solved problems that arises during reinforcement-learning (RL) post-training.

The authors formalize the length-scaling tax (LST) as the percentage increase in response length on queries the policy already answers correctly after RL training.

Why this paper

From ByteDance Seed and 2 others · Part of Agent Harness Optimization, now 82 papers

In one line

Length Self-Distillation reduces excess response length on already-solved queries without sacrificing accuracy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.