Mixture-of-Experts Specialization
Methods for preventing expert collapse and promoting functional specialization in mixture-of-experts models.
4 papers
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.
How this thread developed
Oct 2026 · Bar-Ilan University, NVIDIA
Supervising token-level loss improves routing in sparse mixture-of-experts models
Token-level supervision aligns routing decisions with next-token loss.
1 further paper
Oct 2026 · Peking University, Tencent Hunyuan
Routing-signature diversity prevents expert collapse and boosts MoE specialization
Proposes Distributional Orthogonalization Loss to directly record and penalize expert token-overlap, preventing collapse in MoE models.
Oct 2026 · Indian Institute of Technology Bombay, IBM Research
Soft routing among process experts helps PDE forecasting, but gains prove seed-dependent
Uses soft routing among process-biased experts for PDE forecasting, with seed-dependent gains.
3 of 4 papers shown