Masked co-adaptive fine-tuning cuts expert fetches by over 23% for MoE models.

The method jointly trains routers and experts under learnable binary masks, reducing time per output token by up to 16% in offloaded inference.

Big Tech
Junfeng Wu · Zehao Fan · Hadjer Benmeziane · Kaoutar El Maghraoui · Liu Liu · Yinan Wang

Rensselaer Polytechnic Institute · IBM Research

Research Digest··3 min read
Wu et al.

The authors introduce MaskCoFT, a fine-tuning approach for mixture-of-experts (MoE) language models that trains both routers and experts using only the cross-entropy loss.

Why this paper

From IBM Research and Rensselaer Polytechnic Institute

In one line

MaskCoFT cuts expert fetches per token by 23.7% on Mixtral-8x7B and 10.1% on DeepSeek-V2-Lite while staying above base model accuracy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.