Mixture-of-experts models route each token to a small subset of specialized parameter blocks, allowing high total capacity while activating relatively few parameters.
Reusing Earlier Expert Choices Improves Sparse Model Training Efficiency
HERO-MoE feeds previous layers’ routing distributions into later routers, lowering training loss with less than 1 percent additional memory and computation.
Chinese Tech
Junxiang Qiu · Zhengsu Chen · Xinting Hu · Shuo Wang · Hengheng Zhang · Shaofeng Zhang · +3 more
University of Science and Technology of China · Huawei Inc. · Southeast University
Research Digest··2 min read
Qiu and colleagues modify mixture-of-experts routing so each layer can consult the expert-selection distributions produced for the same token by preceding layers.
Why this paper
From Huawei Inc. and 2 others
In one line
HERO-MoE reuses preceding routing distributions to improve MoE router quality, reducing training loss with minimal overhead.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 8B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§