The model is divided into a dense front end and a mixture-of-experts section.
Segment-level expert routing gives one model four inference budgets
Stepped MoE selects and periodically refreshes expert subnetworks, supporting 1 billion to 4 billion active-parameter configurations without separate pruning models.
Big Tech
Arnav Kundu · Zhaoyang Xu · Bairu Hou · Chang Gao · Reed Li · Tao Lei
Apple
Research Digest··2 min read
Kundu et al.
Why this paper
From Apple
In one line
Stepped MoE combines configurable active parameter budgets with segment-level expert routing, outperforming comparable dense models while retaining similar latency on memory-constrained devices.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§