Leverage score sampling offers scalable diversity for language model pretraining data selection

The method selects samples that expand determinantal volume, achieving up to 72× speedup and improved downstream performance across seven tasks.

Top University
Zailin Ma · Quzhe Huang · Yujun Li · Congyuan Rao · Yaodong Yang

Peking University · Wizard Intelligence Learning Lab · Tsinghua University

Research Digest··2 min read
Ma et al.

The authors propose Leverage Score Sampling (Lev) for selecting diverse pretraining data.

Why this paper

From Peking University and 2 others

In one line

Leverage Score Sampling selects pretraining data to maximize geometric diversity, achieving up to 72x speedup and better downstream accuracy than prior methods.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.