The authors evaluated long-document question answering pipelines on MMLongBench-Doc and LongDocURL.
Agentic document QA pays off mainly with stronger vision-language models
Across two benchmarks, agent benefits depended on reader capability, while image retrieval reduced token use and generally outperformed text retrieval.
Big Tech
Kenan E. Ak · Jay Mohta · Gwang Gook Lee · Yan Xu · Dimitrios Dimitriadis
Amazon.com
Research Digest··2 min read
The authors compare static and agentic pipelines for answering questions about long, visually complex documents using frontier and open-weight vision-language models.
Why this paper
From Amazon.com · Part of Context Engineering for Agents, now 30 papers
In one line
In long-document VLM QA, image retrieval dominates text retrieval in accuracy and token efficiency, and no single pipeline works best for all questions.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§