DolphinBench measures agent memory through tasks, cost, and latency

The benchmark tests whether agents can use long user histories to complete knowledge-work tasks while reporting practical efficiency trade-offs.

Industry
Soumil Rathi · Deshraj Yadav · Taranjeet Singh

Mem0

Research Digest··2 min read
Rathi, Yadav, and Singh introduce a benchmark that evaluates long-term memory through agent task completion rather than prompted question answering.

5 million tokens of user messages.

Why this paper

From Mem0 · Part of Memory Management for Agents, now 37 papers

In one line

DolphinBench evaluates agent memory using action-based tasks, cost/latency metrics, and verified tests.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.