Dataset turns failed agent traces into diagnosis and repair supervision

The authors compile 50,228 diagnosed failures and show that proposed corrections often outperform retrying the original action.

Industry
Kunlun Zhu · Xuyan Ye · Yibo Li · Cheng Qian · Beibin Li · Heng Ji

Apodex

Research Digest··2 min read
Zhu et al.

The authors collected 50,228 error-diagnosis pairs from 9,961 tasks spanning 33 text-based environments, 19 agent harness families, and 23 policy models.

Why this paper

From Apodex

In one line

A dataset of 50,228 error-diagnosis pairs enables LLM agents to learn from failures, improving correction success and diagnosis accuracy.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.