Nahshan et al.
Supervising token-level loss improves routing in sparse mixture-of-experts models
Two new mechanisms align expert selection with next-token prediction error, boosting accuracy on question-answering benchmarks.
Big Tech
Yury Nahshan · Nati Daniel · Jacob Goldberger · Yoli Shavit
Bar-Ilan University · NVIDIA
Research Digest··3 min read
The authors introduce token-error supervision (TES) and affinity-concentration supervision (ACS), two methods that directly align sparse routing decisions in mixture-of-experts language models with the realized next-token cross-entropy loss.
Why this paper
From NVIDIA and Bar-Ilan University
In one line
Sparse MoE routing improves when token-level cross-entropy loss directly supervises expert selection.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§