The authors trained byte Transformers on text data using token-superposition training and hash embeddings, comparing their scaling behavior to subword Transformers.
Byte Transformers outperform subword models at scale and learn implicit token-like abstractions
Using token-superposition training and hash embeddings, byte models consistently beat subword counterparts, and they develop segmentation-like positions for local context, enabling efficient speculative decoding.
Top University
Jie Wang · Shiwei Luo · Qi Zhang · Yuanbin Wu
East China Normal University · Fudan University
Research Digest··2 min read
Jie Wang, Shiwei Luo, Qi Zhang, and Yuanbin Wu train flat byte Transformers without specialized tokenization architectures and show that they outperform subword Transformers as model size scales.
Why this paper
From Fudan University and East China Normal University
In one line
Byte Transformers with token-superposition training and hash embeddings outperform subword Transformers while building internal text abstractions.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§