The authors propose Leverage Score Sampling (Lev) for selecting diverse pretraining data.
Leverage score sampling offers scalable diversity for language model pretraining data selection
The method selects samples that expand determinantal volume, achieving up to 72× speedup and improved downstream performance across seven tasks.
Top University
Zailin Ma · Quzhe Huang · Yujun Li · Congyuan Rao · Yaodong Yang
Peking University · Wizard Intelligence Learning Lab · Tsinghua University
Research Digest··2 min read
Ma et al.
Why this paper
From Peking University and 2 others
In one line
Leverage Score Sampling selects pretraining data to maximize geometric diversity, achieving up to 72x speedup and better downstream accuracy than prior methods.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§