The authors begin with a pretrained autoregressive teacher, which normally generates a sequence one token at a time.
Distilled language models generate sequences through one-step continuous flows
Gumbel Straight Flow uses a pretrained autoregressive model to define direct, non-intersecting paths from noise to token sequences.
Big Tech
Yeongmin Kim · Arnaud Doucet · Andrew Campbell · Valentin De Bortoli · Thomas Mensink · David Ruhe
Google DeepMind Amsterdam
Research Digest··3 min read
Kim and colleagues introduce Gumbel Straight Flow (GSF), a method for distilling sequential autoregressive generation into a continuous map that can produce tokens in parallel.
Why this paper
From Google DeepMind Amsterdam
In one line
Gumbel Straight Flow distills autoregressive models into one-step flow maps that outperform few-step baselines.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§