The authors built 50 healthcare and 80 finance tasks around brief, realistic requests accompanied by workspaces containing documents such as PDFs, spreadsheets and Word files.
Top AI agents complete fewer than one-quarter of professional tasks
DAYJOB tests whether agents can interpret underspecified requests, inspect workplace files and produce fully correct healthcare and finance deliverables.
Industry
Stephanie Finley · Liudas Panavas · Thomas Mikkelson · Cam Hinton · Stacey Ganss · Bradley Monton · +9 more
Surge AI
Research Digest··2 min read
Finley and colleagues introduce DAYJOB, a benchmark of 130 long-horizon assignments designed by professionals in healthcare and finance.
Why this paper
From Surge AI
In one line
Current AI agents fail most long-horizon professional tasks, with the best passing only about a quarter.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§