Token-level teacher routing beats prompt-level routing in multi-teacher distillation

A plug-in framework called MOPD-Router weights supervision from the full teacher pool at each token, using a specialization-aware metric that improves both unlabeled and domain-labeled distillation.

Chinese Tech
Tianze Xu · Yanzhao Zheng · Zhentao Zhang · Yuanqiang Yu · Chao Ma · Jihuai Zhu · +7 more

Shanghai Jiao Tong University · GAIR · Alibaba Group · University of Science and Technology of China · Shanghai Innovation Institute

Research Digest··2 min read
The authors propose MOPD-Router, a framework for multi-teacher on-policy distillation that replaces prompt-level hard routing with token-level routing across all teachers, requiring no domain labels and no separate routing model.

The authors study multi-teacher on-policy distillation (MOPD), where a student generates rollouts and multiple specialist teachers provide dense token-level supervision.

Why this paper

From Alibaba Group and 4 others · Released code

In one line

Token-level ExpertAlign routing combines specialized teachers without domain labels and outperforms prompt-level routing and mean aggregation across four distillation settings.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.