Top AI agents complete fewer than one-quarter of professional tasks

DAYJOB tests whether agents can interpret underspecified requests, inspect workplace files and produce fully correct healthcare and finance deliverables.

Industry
Stephanie Finley · Liudas Panavas · Thomas Mikkelson · Cam Hinton · Stacey Ganss · Bradley Monton · +9 more

Surge AI

Research Digest··2 min read
Finley and colleagues introduce DAYJOB, a benchmark of 130 long-horizon assignments designed by professionals in healthcare and finance.

The authors built 50 healthcare and 80 finance tasks around brief, realistic requests accompanied by workspaces containing documents such as PDFs, spreadsheets and Word files.

Why this paper

From Surge AI

In one line

Current AI agents fail most long-horizon professional tasks, with the best passing only about a quarter.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.