The authors designed a hierarchical router with a binary tree of depth log₂ E, where E=16 experts are leaf nodes.
Tree-based routing balances expert use without auxiliary loss in MoE models
The BRANCH-MoE architecture uses exponential moving average anchors at binary tree nodes to guarantee utilization and reduce communication.
Big Tech
Gang Fu · Adel Javanmard · MohammadHossein Bateni · Vahab Mirrokni
Google Research · University of Southern California
Research Digest··3 min read
The authors propose BRANCH-MoE, a routing architecture that places experts at leaves of a binary decision tree.
Why this paper
From Google Research and University of Southern California
In one line
BRANCH-MoE uses a binary decision tree with EMA anchors to balance expert utilization without auxiliary losses.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§