The authors propose LCT, a tokenizer that decouples the discovery of latent structural units from the construction of the final vocabulary.
Morphology-aware tokenizer beats compression-only baselines across 104 languages
The Latent Core Tokenizer separates structural discovery from vocabulary construction, improving both tokenization efficiency and downstream performance without increasing cross-lingual disparity.
Big Tech
Felermino D. M. A. Ali · Millicent Ochieng · Ogbemi Ekwejunor-Etchie · Ade Famoti · Jacki O'Neill · Debjit Paul
Microsoft Research Africa · Microsoft Research Accelerator · Microsoft Research India
Research Digest··2 min read
Ali et al.
Why this paper
From Microsoft Research Africa and 2 others
In one line
Latent Core Tokenizer improves multilingual tokenization by separating structural discovery from vocabulary construction.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§