The authors constructed DSV-Mem (Dense Stateful Visual Memory) to evaluate conversational multimodal large language model (MLLM) agents in professional contexts such as AI research, engineering design, and product management.
New benchmark tests MLLM agents on dense, stateful professional visuals
DSV-Mem reveals that frontier multimodal LLMs score below 45%, with state evolution proving the dominant challenge.
Big Tech
Jike Zhong · Ritwick Chaudhry · Xuanbai Chen · Tianchen Zhao · Linghan Xu · Yifan Xing · +1 more
Amazon AGI · University of Southern California
Research Digest··2 min read
Zhong et al.
Why this paper
From Amazon AGI and University of Southern California
In one line
The best MLLM agent scores below 45% on the DSV-Mem benchmark for professional multimodal memory.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§