agentic evaluation
Measuring AI agents' performance in complex tasks, highlighting failures in planning, monitoring, and real
- Papers
- 25
- Released code
- 7
- First seen
- June 2026
- Latest
- Sept 2026
23 papers in the last two months, against 2 in the two before.
Who is working on it
The papers
Most central to this idea first, not most recent.
- Big Techcs.AI
Replayable news timelines test how LLM agents update forecasts
Georgia Institute of Technology, Amazon · Sept 2026
- Industrycs.CRcode
Cyber agents hide critical weaknesses behind successful attack workflows
National Research Council, University of Windsor · Sept 2026
- Independentcs.AI
Competitive games provide a dynamic benchmark for language model strategy
Sept 2026
- Top Universitycs.AI
Cross-device workflows expose major weaknesses in today’s GUI agents
Beihang University, Beijing Institute of Technology · Sept 2026
- Big Techcs.AIcode
Civilization benchmark exposes agents’ weak monitoring and long-term follow-through
University of Oxford, Tony Blair Institute for Global Change · Sept 2026
- Top Universitycs.CVcode
MLLMs fail to sustain goal-directed navigation in real-scale 3D city
Shanghai Jiao Tong University, National University of Singapore · Aug 2026
- Industrycs.LG
Domain context improves LLM agents’ time-series root-cause attribution
ETH Zürich, CSEM SA · Aug 2026
- Independentcs.AI
Benchmark tests whether coding agents improve governance without breaking repositories
Sept 2026
- AI Startupcs.AIcode
Goal pressure exposes wide gaps in security agents’ scope adherence
dreadnode · Sept 2026
- Independentcs.SE
Refresh schedules, not cache age, determine agent answer staleness
Sept 2026
- Industrycs.IR
Benchmark tests retrieval for agent-written queries across 190 million web pages
Perplexity AI · Sept 2026
- Independentcs.SE
LLMs can build agent harnesses, but improvements transfer poorly
Sept 2026
- Industrycs.CL
Tool-call traces expose extraction failures that source fidelity misses
Infineon Technologies AG · Sept 2026
- Chinese Techcs.AIcode
Models remain weak at steering coding agents through long tasks
DreamX Team, Alibaba Group · Sept 2026
- Top Universitycs.CL
LLM agents struggle to infer wellbeing from long-term wearable data
Dartmouth College, University of Cambridge · Aug 2026
- Chinese Techcs.AI
Process-level tests expose distinct mathematical agent capabilities in LLMs
Sun Yat-sen University, Tencent Youtu Lab · Aug 2026
- Independentcs.SE
Enterprise agents struggle when facts are hidden across business systems
Sept 2026
- Top Universitycs.LGcode
New benchmark standardizes evaluation of coding agent harnesses
TokenRhythm Technologies, Infinigence AI · June 2026
- Independentcs.AI
Deterministic guardrails expose gains that fool LLM-based judges
Sept 2026
- Independentcs.CL
LLM agents do not reliably improve from their own experience
Sept 2026
- Independentcs.SEcode
Coding agent benchmarks miss review compliance gaps
Sept 2026
- Top Universitycs.CL
Simulated user personas improve agents’ handling of multi-turn workflows
State Key Laboratory of Multimedia Information Processing, Peking University · Sept 2026
- Big Techcs.AI
Structured behavioral abstractions improve diagnosis of failures in LLM agents
Tsinghua University, Microsoft Research · Sept 2026
- Industrycs.AI
Dual-layer knowledge graph connects fragmented pharmaceutical process-development documents
Sanofi US · Sept 2026
- Top Universitycs.AI
New benchmark reveals coordination as distinct bottleneck for LLM agents
University of Edinburgh, University of Oxford · June 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.