LAM combines three components: a deterministic deduplication rule with a bound on how much substitution can perturb retrieval scores, a memory manager that preserves the cached prefix and overlaps compaction with inference, and a performance model for estimating compaction costs.
LAM Cuts Agent Memory While Bounding Retrieval Score Errors
A deterministic compaction method removed redundant observations before model prefill while preserving nearly all measured evidence across 600 agent trajectories.
Big Tech
Baixi Sun · Le Chen · Anjir Ahmed Chowdhury · Xiaolong Ma · Chih-Hsuan Yang · Mingze Xia · +7 more
Argonne National Laboratory · University of Houston · Oregon State University · IBM Research
Research Digest··2 min read
The authors introduce LAM, a memory framework that deduplicates agent observations without using an LLM summarizer, preserves cached context prefixes, and estimates compaction costs before deployment.
Why this paper
From IBM Research and 3 others
In one line
LAM trims 22.47% of agent memory tokens with 99.984% evidence retention and predicts up to 91.6x speedup by scheduling compaction before prefill.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§