One nested language model works across every tested layer depth

Stochastic supervision made all 20 layer prefixes usable while preserving full-capacity performance.

Big Tech
Zhilin Guo · Boqiao Zhang · Hakan Aktas · Kyle Fogarty · Nursena Koprucu Aslan · Wenzhao Li · +11 more

University of Cambridge · University of British Columbia · Google

Research Digest··2 min read
Guo and colleagues trained a 200 million-parameter Transformer whose computation can stop after any layer, yielding a continuum of smaller language models from one artifact.

The authors trained on 20 billion FineWeb-Edu tokens, using the same data stream across methods.

Why this paper

From Google and 2 others

In one line

Training nested-capacity Transformers with stochastic prefix supervision yields valid language models at every depth from one run.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 200M)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.