Factorized n-gram memory with basis-level gating improves language model efficiency

FactorEngram retrieves sparse coefficients over a shared dictionary of basis vectors, enabling contextual modulation of each memory component individually.

Chinese Tech
Bowen Yang · Jingbo Zhou · Qinghong Miao · Hua Wu

Baidu Inc. · Nanyang Technological University

Research Digest··2 min read
The authors propose FactorEngram, a memory module for language models that replaces monolithic n-gram embeddings with sparse coefficients over a shared dictionary.

FactorEngram augments a Transformer with an auxiliary memory branch that retrieves sparse coefficients from hash tables, one per n-gram pattern (single tokens and multi-token n-grams).

Why this paper

From Baidu Inc. and Nanyang Technological University

In one line

FactorEngram improves language models by factorizing n-gram memory into sparse, context-gated basis vectors.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.