AutoSciBench automatically generates scientific agent benchmarks that adapt as agents improve.

The framework iteratively refines tasks to eliminate shortcuts and boost difficulty, outperforming human-curated benchmarks in three scientific domains.

Top University
Dongki Kim · Namkyeong Lee · Surag Nair · Carl Edwards · Xiner Li · Edward De Brouwer · +4 more

KAIST · Genentech

Research Digest··2 min read
Kaplan et al.

The authors developed AutoSciBench, which organizes benchmark task construction around two levels: a high-level concept (specifying scientific domain, data modality, and reasoning approach) and a low-level recipe (detailing how the question, environment, and ground truth are built).

Why this paper

From KAIST and Genentech

In one line

AutoSciBench generates and refines scientific-agent benchmarks, producing harder and higher quality tasks than human-curated ones.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.