AI agents reproduce only 41% of machine learning papers with full code and weights provided

RECLAIM benchmark evaluates agents on 100 NeurIPS 2025 papers across three difficulty tiers, revealing steep drops in success when agents must retrain or reimplement

Academic
Mithil Salunkhe · Haochen Ding · Samridhi Verma · Volodymyr Kindratenko

University of Illinois Urbana-Champaign · National Center for Supercomputing Applications

Research Digest··3 min read
The authors introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers with predefined reproduction targets and GPU-hour budgets.

The authors constructed RECLAIM, a benchmark of 100 papers from NeurIPS 2025, each with a fixed target result, success criterion, and GPU-hour budget.

Why this paper

From University of Illinois Urbana-Champaign and National Center for Supercomputing Applications · Part of Repository-Level Agent Benchmarks, now 3 papers

In one line

AI agents reproduce ML paper results at 15-41% rates depending on code availability; most failures lack verification against the paper's numbers.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.