Mixture-of-Experts Inference Efficiency
Methods for cutting compute and memory costs when serving mixture-of-experts models, including efficient routers, expert pruning, and batched decoding.
3 papers
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.