Byte Transformers outperform subword models at scale and learn implicit token-like abstractions

Using token-superposition training and hash embeddings, byte models consistently beat subword counterparts, and they develop segmentation-like positions for local context, enabling efficient speculative decoding.

Top University
Jie Wang · Shiwei Luo · Qi Zhang · Yuanbin Wu

East China Normal University · Fudan University

Research Digest··2 min read
Jie Wang, Shiwei Luo, Qi Zhang, and Yuanbin Wu train flat byte Transformers without specialized tokenization architectures and show that they outperform subword Transformers as model size scales.

The authors trained byte Transformers on text data using token-superposition training and hash embeddings, comparing their scaling behavior to subword Transformers.

Why this paper

From Fudan University and East China Normal University

In one line

Byte Transformers with token-superposition training and hash embeddings outperform subword Transformers while building internal text abstractions.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.