mixture of experts
A neural architecture with many specialized sub-networks, only a few activated per input via a router, balancing capacity and efficiency.
- Papers
- 6
- Released code
- 1
- First seen
- July 2026
- Latest
- Sept 2026
3 papers in the last two months, against 3 in the two before.
The papers
Most central to this idea first, not most recent.
- Top Universitycs.LG
Sparse mixture-of-experts models overfit repeated training data faster
Stanford University, Paul G. Allen School of Computer Science, University of Washington · Sept 2026
- Independentcs.CLcode
A 2.8-trillion-parameter open MoE model approaches frontier performance.
July 2026
- Top Universitycs.LG
A framework predicts RL post-training outcomes without running reinforcement learning
MIT CSAIL · Sept 2026
- Big Techcs.CV
World state registers enable consistent multi-agent video generation across views
University of California, Los Angeles, Adobe Research · July 2026
- Chinese Techcs.LG
Shared-prefix training accelerates reinforcement learning for hybrid-attention agents
Ant Group · Sept 2026
- Big Techcs.LG
A lightweight PyTorch-native framework matches Megatron-based agentic RL training performance.
NVIDIA · July 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.