The authors equipped a general-purpose vision-language model with tools for directly observing an interactive environment, retaining its complete visual history, and selecting previous observations for renewed inspection.
Visual memory harness helps multimodal agents solve long interactive tasks
VISTA lets a general-purpose model retrieve and reorganize original visual observations while reasoning across games and puzzles.
Academic
Qiushi Han · Keya Hu · Linlu Qiu · Cathy Wu · Kaiming He
Massachusetts Institute of Technology
Research Digest··2 min read
Han et al.
Why this paper
From Massachusetts Institute of Technology
In one line
VISTA uses lossless visual memory to boost multimodal model performance to 100% on ARC-AGI-3.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§