The authors manually created and independently validated 423 questions spanning 13 languages, eight evidence modalities and 13 primary domains.
Web agents struggle with multilingual, multimodal evidence hunts across the web
Across 423 difficult questions in 13 languages, the best evaluated browsing configuration answered fewer than one-third correctly.
Chinese Tech
Alham Fikri Aji · Faiz Rizki Ramadhan · Zayd M. K. Zuhri · Seung Hun Eddie Han · Ryandito Diandaru · Qinrong Cui · +11 more
Mohamed bin Zayed University of Artificial Intelligence · Mila – Quebec Artificial Intelligence Institute · Inception AI · Alibaba Group · AI Singapore
Research Digest··3 min read
Aji and colleagues introduce HyperBrowseComp, a benchmark designed to test whether web agents can pursue implicit clues across languages and source types, rather than retrieve straightforward facts.
Why this paper
From Alibaba Group and 4 others
In one line
HyperBrowseComp is a 423 question benchmark across 13 languages and multiple modalities that challenges web browsing agents.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§