Visual memory harness helps multimodal agents solve long interactive tasks

VISTA lets a general-purpose model retrieve and reorganize original visual observations while reasoning across games and puzzles.

Academic
Qiushi Han · Keya Hu · Linlu Qiu · Cathy Wu · Kaiming He

Massachusetts Institute of Technology

Research Digest··2 min read
Han et al.

The authors equipped a general-purpose vision-language model with tools for directly observing an interactive environment, retaining its complete visual history, and selecting previous observations for renewed inspection.

Why this paper

From Massachusetts Institute of Technology

In one line

VISTA uses lossless visual memory to boost multimodal model performance to 100% on ARC-AGI-3.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.