Agentic document QA pays off mainly with stronger vision-language models

Across two benchmarks, agent benefits depended on reader capability, while image retrieval reduced token use and generally outperformed text retrieval.

Big Tech
Kenan E. Ak · Jay Mohta · Gwang Gook Lee · Yan Xu · Dimitrios Dimitriadis

Amazon.com

Research Digest··2 min read
The authors compare static and agentic pipelines for answering questions about long, visually complex documents using frontier and open-weight vision-language models.

The authors evaluated long-document question answering pipelines on MMLongBench-Doc and LongDocURL.

Why this paper

From Amazon.com · Part of Context Engineering for Agents, now 30 papers

In one line

In long-document VLM QA, image retrieval dominates text retrieval in accuracy and token efficiency, and no single pipeline works best for all questions.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.