The authors created a benchmark spanning four tasks from astrophysics and genomics.
Agent performance depends more on system configuration than model size
Across four scientific tasks and more than 18,000 trajectories, information, runtime, tools, and model choice jointly shaped agent capability and reliability.
Research Lab
Luis Wiedmann · Leander Girrbach · Cordelia Schmid · Zeynep Akata
Technical University of Munich · MCML · Helmholtz Munich · Inria · École normale supérieure
Research Digest··2 min read
Wiedmann and colleagues evaluate coding agents as configurable systems rather than treating the underlying language model as the sole unit of analysis.
Why this paper
From Helmholtz Munich and 6 others · Released code
In one line
Agent performance depends more on system configuration, especially supplied information, than on model choice alone, and reliable verification requires system tools rather than prompts.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§