Web agents struggle with multilingual, multimodal evidence hunts across the web

Across 423 difficult questions in 13 languages, the best evaluated browsing configuration answered fewer than one-third correctly.

Chinese Tech
Alham Fikri Aji · Faiz Rizki Ramadhan · Zayd M. K. Zuhri · Seung Hun Eddie Han · Ryandito Diandaru · Qinrong Cui · +11 more

Mohamed bin Zayed University of Artificial Intelligence · Mila – Quebec Artificial Intelligence Institute · Inception AI · Alibaba Group · AI Singapore

Research Digest··3 min read
Aji and colleagues introduce HyperBrowseComp, a benchmark designed to test whether web agents can pursue implicit clues across languages and source types, rather than retrieve straightforward facts.

The authors manually created and independently validated 423 questions spanning 13 languages, eight evidence modalities and 13 primary domains.

Why this paper

From Alibaba Group and 4 others

In one line

HyperBrowseComp is a 423 question benchmark across 13 languages and multiple modalities that challenges web browsing agents.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.