Deep research agents struggle with long, multimodal evidence chains

Mr.LHDR tests whether agents can integrate consequential visual and textual evidence across deeply dependent research steps.

PaperTop Universitycs.AIarXiv:2609.11318v1
Minghao Guo · Meng Cao · Sui Zhao · Siyu Ning · Xin Wang · Haoze Zhao · +7 more

Mohamed bin Zayed University of Artificial Intelligence · University of Science and Technology of China · Zhejiang University

Research Digest··2 min read
The authors introduce a benchmark whose questions require an average of 12.1 intermediate conclusions and include non-text evidence that affects the reasoning process. The strongest evaluated system reached 43.1% Overall Accuracy but only 34.3% Strict Accuracy, exposing failures that final-answer scoring can obscure.

What they did

The authors built Mr.LHDR around real-world research questions spanning eight categories. Each question derives from a hidden node-relation graph encoding the evidence and dependencies needed to reach a short, uniquely verifiable answer; questions average 12.1 necessary intermediate conclusions and a dependency depth of 10.4.

Every task includes at least one consequential non-text element, such as an image, map, PDF, logo, chart, table, or video frame. General models, deep-research systems, and agent frameworks were assessed using Overall Accuracy, Strict Accuracy, Checklist Score, and a Dependency-Aware Checklist Score that accounts for whether intermediate conclusions respect their required dependencies.

Key findings

  • The strongest system achieved 43.1% Overall Accuracy and 34.3% Strict Accuracy.
  • The gap between these metrics indicates that correct final answers can conceal incomplete or inconsistent research trajectories.
  • Removing images lowered the Dependency-Aware Checklist Score by 12.6 points, showing that visual evidence materially affected task completion.
  • Strict Accuracy consistently decreased as the required reasoning chains grew longer.

Why it matters

Mr.LHDR shifts evaluation from isolated retrieval or medium-length browsing toward sustained evidence integration. By scoring intermediate conclusions and their dependencies, it can distinguish agents that merely arrive at the right answer from those that execute a coherent, auditable research process.

Caveats

The reported results concern one benchmark with graph-structured, uniquely verifiable questions and may not capture open-ended research where evidence is ambiguous or multiple answers are defensible. The abstract also does not specify system-by-system results, annotation reliability, contamination controls, or how performance transfers to unconstrained production environments.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.