The author built a streaming-recall benchmark that crosses retention rules with selection rules in a factorial design.
Memory Tests Must Separate What Agents Keep From What They Retrieve
A factorial benchmark shows that changing an agent’s access to history can dwarf the measured benefit of query-aware retrieval.
Top University
Juli Huang
Stanford University
Research Digest··3 min read
Thread:Memory Management for Agents
Huang evaluates retention, which determines what remains in bounded memory, separately from selection, which chooses stored information after a query arrives.
Why this paper
From Stanford University · Part of Memory Management for Agents, now 43 papers
In one line
Memory evaluations confound retention with retrieval; most apparent query-awareness gains are actually history access, so benchmarks should hold access fixed and report both separately.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§