What they did
The authors built Mr.LHDR around real-world research questions spanning eight categories. Each question derives from a hidden node-relation graph encoding the evidence and dependencies needed to reach a short, uniquely verifiable answer; questions average 12.1 necessary intermediate conclusions and a dependency depth of 10.4.
Every task includes at least one consequential non-text element, such as an image, map, PDF, logo, chart, table, or video frame. General models, deep-research systems, and agent frameworks were assessed using Overall Accuracy, Strict Accuracy, Checklist Score, and a Dependency-Aware Checklist Score that accounts for whether intermediate conclusions respect their required dependencies.
Key findings
- The strongest system achieved 43.1% Overall Accuracy and 34.3% Strict Accuracy.
- The gap between these metrics indicates that correct final answers can conceal incomplete or inconsistent research trajectories.
- Removing images lowered the Dependency-Aware Checklist Score by 12.6 points, showing that visual evidence materially affected task completion.
- Strict Accuracy consistently decreased as the required reasoning chains grew longer.
Why it matters
Mr.LHDR shifts evaluation from isolated retrieval or medium-length browsing toward sustained evidence integration. By scoring intermediate conclusions and their dependencies, it can distinguish agents that merely arrive at the right answer from those that execute a coherent, auditable research process.
Caveats
The reported results concern one benchmark with graph-structured, uniquely verifiable questions and may not capture open-ended research where evidence is ambiguous or multiple answers are defensible. The abstract also does not specify system-by-system results, annotation reliability, contamination controls, or how performance transfers to unconstrained production environments.