The authors introduce MaskCoFT, a fine-tuning approach for mixture-of-experts (MoE) language models that trains both routers and experts using only the cross-entropy loss.
Masked co-adaptive fine-tuning cuts expert fetches by over 23% for MoE models.
The method jointly trains routers and experts under learnable binary masks, reducing time per output token by up to 16% in offloaded inference.
Big Tech
Junfeng Wu · Zehao Fan · Hadjer Benmeziane · Kaoutar El Maghraoui · Liu Liu · Yinan Wang
Rensselaer Polytechnic Institute · IBM Research
Research Digest··3 min read
Wu et al.
Why this paper
From IBM Research and Rensselaer Polytechnic Institute
In one line
MaskCoFT cuts expert fetches per token by 23.7% on Mixtral-8x7B and 10.1% on DeepSeek-V2-Lite while staying above base model accuracy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§