The authors examined expert parallelism, where a Mixture-of-Experts model distributes its expert subnetworks across GPUs and routes each token to selected experts.
Host-centric routing makes MoE inference practical on consumer GPUs
CoMoE restructures token dispatch and aggregation around host memory, reducing communication overhead on multi-GPU systems without peer-to-peer links.
Chinese Tech
Ruwen Fan (Jimmy) · Yuezhi Zu (Jimmy) · Junru Li (Jimmy) · Qingda Hu (Jimmy) · Xinjun (Jimmy) · Yang · +2 more
Research Digest··2 min read
Fan et al.
Why this paper
From Alibaba Cloud Computing and Tsinghua University
In one line
CoMoE uses host-centric token routing to make multi-GPU MoE inference on consumer GPUs up to 1.46x faster and substantially cheaper than A800 deployments.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (gpu RTX 5090)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§