The authors developed AutoSciBench, which organizes benchmark task construction around two levels: a high-level concept (specifying scientific domain, data modality, and reasoning approach) and a low-level recipe (detailing how the question, environment, and ground truth are built).
AutoSciBench automatically generates scientific agent benchmarks that adapt as agents improve.
The framework iteratively refines tasks to eliminate shortcuts and boost difficulty, outperforming human-curated benchmarks in three scientific domains.
Top University
Dongki Kim · Namkyeong Lee · Surag Nair · Carl Edwards · Xiner Li · Edward De Brouwer · +4 more
KAIST · Genentech
Research Digest··2 min read
Kaplan et al.
Why this paper
From KAIST and Genentech
In one line
AutoSciBench generates and refines scientific-agent benchmarks, producing harder and higher quality tasks than human-curated ones.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§