The authors formalize the length-scaling tax (LST) as the percentage increase in response length on queries the policy already answers correctly after RL training.
Online distillation cuts needless verbosity in RL-trained models
Length Self-Distillation reduces excess response length on easy queries without sacrificing accuracy on hard ones.
Chinese Tech
Xu Wan · Wenyue Xu · Shengjie Zhao · Mingyang Sun
ByteDance Seed · Tongji University · Peking University
Research Digest··3 min read
Thread:Agent Harness Optimization
The authors identify and quantify the 'length-scaling tax'—excess response length on already-solved problems that arises during reinforcement-learning (RL) post-training.
Why this paper
From ByteDance Seed and 2 others · Part of Agent Harness Optimization, now 82 papers
In one line
Length Self-Distillation reduces excess response length on already-solved queries without sacrificing accuracy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§