The authors trained a persistent textual-memory writer against a fixed downstream reader.
Counterfactual credit teaches language models which memories stay useful
MGPO trains a memory writer by measuring each rewrite’s incremental contribution to present and future task performance.
Top University
Jiaming Tang · Mingyan Liu · Armin Sarabi
University of Michigan
Research Digest··2 min read
Tang, Liu and Sarabi address a central problem in persistent language-model memory: a rewrite may help only much later, and its apparent value may actually come from information already stored.
Why this paper
From University of Michigan
In one line
MGPO isolates incremental value of memory rewrites via counterfactual credit, reducing memory length by 80% while improving extraction.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§