Search agents struggle to find broad evidence for open research problems

INSPIRE evaluates literature-search agents on 476 solution-redacted computer science problems and traces failures across evidence exposure, selection and ranking.

Big Tech
Jianrong Ding · Zhengyan Shi · Jianyuan Zhong · Kai Qiu · Qi Dai · Yifan Yang · +2 more

The Chinese University of Hong Kong · Microsoft Research Asia

Research Digest··2 min read
Ding et al.

The authors built 476 computer science search tasks from later research papers.

Why this paper

From Microsoft Research Asia and The Chinese University of Hong Kong

In one line

INSPIRE benchmarks agents searching scientific literature for open problems, achieving 0.284 nDCG@10.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.