Shared memory improves looped Transformers while shrinking context caches

Models trained to reuse the first recursion’s memory achieved better language modeling quality than conventional looped models despite retaining far fewer key-value entries.

Big Tech
Giovanni Monea · Keshav Ramji · Yousef El-Kurdi · Luis A. Lastras · Yoav Artzi · Nathan Godey · +1 more

IBM Research · Cornell University

Research Digest··3 min read
Monea and colleagues pretrained looped language models in which only the first pass writes a full key-value cache, while subsequent passes reuse it and retain only a short local window.

Looped Transformers apply the same layers repeatedly to each token, increasing computation without increasing parameter count.

Why this paper

From IBM Research and Cornell University

In one line

Shared memory across recursions improves quality and reduces memory in looped transformers.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.