VAMR runs one ReAct-style loop over the complete question set.
Video agent answers questions about one video in a single shared reasoning pass
VAMR coordinates tool use and memory across the full question set, beating per-question pipelines on three benchmarks while using fewer reasoning rounds.
Chinese Tech
Runquan Gui · Hanzhu Chen · Zehao Wang · Hanxin Zhu · Xin Li · Zhibo Chen
University of Science and Technology of China · Tencent
Research Digest··2 min read
The authors introduce VAMR, a video agent that answers multiple questions about one recording in a shared tool-use trajectory instead of restarting exploration for each question.
Why this paper
From Tencent and University of Science and Technology of China
In one line
Coordinating multiple questions about one video in a single shared agent trajectory improves accuracy and cuts reasoning rounds and frames.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§