Memory Tests Must Separate What Agents Keep From What They Retrieve

A factorial benchmark shows that changing an agent’s access to history can dwarf the measured benefit of query-aware retrieval.

Top University
Juli Huang

Stanford University

Research Digest··3 min read
Huang evaluates retention, which determines what remains in bounded memory, separately from selection, which chooses stored information after a query arrives.

The author built a streaming-recall benchmark that crosses retention rules with selection rules in a factorial design.

Why this paper

From Stanford University · Part of Memory Management for Agents, now 43 papers

In one line

Memory evaluations confound retention with retrieval; most apparent query-awareness gains are actually history access, so benchmarks should hold access fixed and report both separately.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.