Reusing Earlier Expert Choices Improves Sparse Model Training Efficiency

HERO-MoE feeds previous layers’ routing distributions into later routers, lowering training loss with less than 1 percent additional memory and computation.

Chinese Tech
Junxiang Qiu · Zhengsu Chen · Xinting Hu · Shuo Wang · Hengheng Zhang · Shaofeng Zhang · +3 more

University of Science and Technology of China · Huawei Inc. · Southeast University

Research Digest··2 min read
Qiu and colleagues modify mixture-of-experts routing so each layer can consult the expert-selection distributions produced for the same token by preceding layers.

Mixture-of-experts models route each token to a small subset of specialized parameter blocks, allowing high total capacity while activating relatively few parameters.

Why this paper

From Huawei Inc. and 2 others

In one line

HERO-MoE reuses preceding routing distributions to improve MoE router quality, reducing training loss with minimal overhead.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 8B)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.