← All concepts

agentic evaluation

Measuring AI agents' performance in complex tasks, highlighting failures in planning, monitoring, and real

Papers
25
Released code
7
First seen
June 2026
Latest
Sept 2026

23 papers in the last two months, against 2 in the two before.

Who is working on it

University of Oxford 2

The papers

Most central to this idea first, not most recent.

Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.