The authors compare strict and full-coverage cross-tokenizer on-policy distillation (OPD).
Supervision reliability trumps alignment coverage in cross-tokenizer distillation
Strict token alignment already captures most learning signal; adding span-level supervision reduces accuracy due to weakly aligned gradients.
Chinese Tech
Bingxi Hou · Guochao Jiang · Guofeng Quan · Weiqing Li · Wenfeng Feng · Guohua Liu · +1 more
Alibaba Cloud Computing
Research Digest··3 min read
The authors investigate cross-tokenizer on-policy distillation (OPD) for language models.
Why this paper
From Alibaba Cloud Computing
In one line
In cross-tokenizer on-policy distillation, expanding alignment coverage with span supervision reduces accuracy; compact strict supervision is more effective.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§