The authors split Transformers into modules and stopped gradients at module boundaries.
Shared readouts bring local learning close to backpropagation at billion scale
SOLO trained Transformers with up to 2 billion parameters while reducing activation memory and improving measured pipeline throughput.
Research Lab
Bojian Yin · Shurong Wang · Yuqi Pan · Guoqi Li
Institute of Automation, Chinese Academy of Sciences
Research Digest··2 min read
Yin and colleagues tested a local learning method that lets network modules update independently rather than waiting for gradients to traverse the entire model.
Why this paper
From Institute of Automation, Chinese Academy of Sciences
In one line
SOLO pretrains billion-parameter language models with local learning, achieving accuracy near backpropagation and improving throughput and memory efficiency.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§