process-level evaluation
A granular assessment of a system's intermediate steps or subgoals, rather than just its final output, to diagnose specific failures.
- Papers
- 10
- Released code
- 2
- First seen
- Aug 2026
- Latest
- Sept 2026
10 papers in the last two months, against 0 in the two before.
Who is working on it
The papers
Most central to this idea first, not most recent.
- AI Startupcs.SE
Coding agents often overstate how thoroughly they reviewed files
Mila – Quebec AI Institute, Tara Research · Sept 2026
- Chinese Techcs.AI
Checkpoint testing reveals hidden weaknesses in self-evolving agent memories
Zhejiang University, Alibaba Group · Sept 2026
- Big Techcs.CL
Process-based evaluation reveals where computer-use agents go wrong
NVIDIA · Sept 2026
- Chinese Techcs.AI
Process-level tests expose distinct mathematical agent capabilities in LLMs
Sun Yat-sen University, Tencent Youtu Lab · Aug 2026
- Top Universitycs.AIcode
Deep research agents struggle with long, multimodal evidence chains
Mohamed bin Zayed University of Artificial Intelligence, University of Science and Technology of China · Sept 2026
- Big Techcs.AIcode
Stateful retrieval finds relevant troubleshooting cases throughout support workflows
Microsoft · Sept 2026
- Industrycs.AI
Formal logic solvers can audit and improve language-model reasoning chains
University of Bristol, King Abdullah University of Science and Technology · Sept 2026
- Top Universitycs.AI
Privileged supervision improves action-level credit for language-model agents
Zhejiang University · Sept 2026
- Chinese Techcs.CL
Selecting evidence before summarizing improves long-context reasoning efficiency
Peking University, Baidu Inc. · Sept 2026
- AI Startupcs.CL
Chain-of-thought reasoning traces are legible but not interpretable
ETH Zurich, MIT · Sept 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.