The authors trained on 20 billion FineWeb-Edu tokens, using the same data stream across methods.
One nested language model works across every tested layer depth
Stochastic supervision made all 20 layer prefixes usable while preserving full-capacity performance.
Big Tech
Zhilin Guo · Boqiao Zhang · Hakan Aktas · Kyle Fogarty · Nursena Koprucu Aslan · Wenzhao Li · +11 more
University of Cambridge · University of British Columbia · Google
Research Digest··2 min read
Guo and colleagues trained a 200 million-parameter Transformer whose computation can stop after any layer, yielding a continuum of smaller language models from one artifact.
Why this paper
From Google and 2 others
In one line
Training nested-capacity Transformers with stochastic prefix supervision yields valid language models at every depth from one run.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 200M)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§