The authors constructed RECLAIM, a benchmark of 100 papers from NeurIPS 2025, each with a fixed target result, success criterion, and GPU-hour budget.
AI agents reproduce only 41% of machine learning papers with full code and weights provided
RECLAIM benchmark evaluates agents on 100 NeurIPS 2025 papers across three difficulty tiers, revealing steep drops in success when agents must retrain or reimplement
Academic
Mithil Salunkhe · Haochen Ding · Samridhi Verma · Volodymyr Kindratenko
University of Illinois Urbana-Champaign · National Center for Supercomputing Applications
Research Digest··3 min read
The authors introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers with predefined reproduction targets and GPU-hour budgets.
Why this paper
From University of Illinois Urbana-Champaign and National Center for Supercomputing Applications · Part of Repository-Level Agent Benchmarks, now 3 papers
In one line
AI agents reproduce ML paper results at 15-41% rates depending on code availability; most failures lack verification against the paper's numbers.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§