Sparse feed-forward activations cut language-model inference costs substantially

Model casting and a cheaper gating design reduce feed-forward computation while largely preserving model quality.

Big Tech
Maria Lomeli · Antoine Groudiev · Matthijs Douze · Loïc Cabannes · Pierre Emmanuel Mazaré · François Fleuret · +4 more

Meta FAIR · École Normale Supérieure – PSL · École Normale Supérieure Paris-Saclay

Research Digest··2 min read
Lomeli and colleagues introduce model casting, a mid-training procedure that makes feed-forward network activations sparse enough for dedicated inference kernels to skip most inactive computation.

The authors replace or modify the activation functions in pretrained language models, then continue training with an adaptive sparsity regularizer.

Why this paper

From Meta FAIR and 2 others

In one line

Model casting and LoPA gating sparsify FFN activations, achieving up to 3.31x GPU speedup with minimal quality loss.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.