The authors analyze stochastic Muon without momentum and derive a conditional loss-neutral boundary of 2ρ_b/η, where η is the learning rate and ρ_b corrects for the coherence of minibatch-induced Muon directions.
Muon Splits Loss Stability From Update Reversal During LLM Pretraining
The authors find that Muon can track a stochastic loss-neutral boundary without exhibiting the update-direction reversal associated with gradient descent.
Chinese Tech
Yanzhe Chen · Qifang Zhao · Xiaoxiao Xu · Fanghui Liu
Shanghai Jiao Tong University · Alibaba Inc.
Research Digest··2 min read
Chen and colleagues examine whether the classical edge-of-stability picture for gradient descent also describes language models trained with Muon, an optimizer that replaces matrix gradients with approximately semi-orthogonal directions.
Why this paper
From Alibaba Inc. and Shanghai Jiao Tong University · Released code
In one line
Muon decouples loss neutrality from update reversal, creating separate boundaries in LLM pretraining.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§