← All threads

Mixture-of-Experts Specialization

Methods for preventing expert collapse and promoting functional specialization in mixture-of-experts models.

4 papers

Where this stands

The written synthesis of this thread is for subscribers. Subscribe.

How this thread developed

  1. Oct 2026 · Bar-Ilan University, NVIDIA

    Supervising token-level loss improves routing in sparse mixture-of-experts models

    Token-level supervision aligns routing decisions with next-token loss.

  2. 1 further paper

    Oct 2026 · Peking University, Tencent Hunyuan

    Routing-signature diversity prevents expert collapse and boosts MoE specialization

    Proposes Distributional Orthogonalization Loss to directly record and penalize expert token-overlap, preventing collapse in MoE models.

  3. Oct 2026 · Indian Institute of Technology Bombay, IBM Research

    Soft routing among process experts helps PDE forecasting, but gains prove seed-dependent

    Uses soft routing among process-biased experts for PDE forecasting, with seed-dependent gains.

3 of 4 papers shown