benchmark
A standard set of tasks or data used to evaluate and compare the performance of AI systems.
- Papers
- 14
- Released code
- 2
- First seen
- June 2026
- Latest
- Sept 2026
12 papers in the last two months, against 2 in the two before.
The papers
Most central to this idea first, not most recent.
- Top Universitycs.AI
API benchmark scores poorly predict performance in consumer chatbot interfaces
Stanford University · Sept 2026
- Industrycs.IR
Benchmark tests retrieval for agent-written queries across 190 million web pages
Perplexity AI · Sept 2026
- Industrycs.CV
Identity preservation remains a distinct challenge for generative image models
Phota Labs Research · Sept 2026
- Independentcs.SE
Enterprise agents struggle when facts are hidden across business systems
Sept 2026
- Independentcs.SE
Refresh schedules, not cache age, determine agent answer staleness
Sept 2026
- Chinese Techcs.AIcode
Models remain weak at steering coding agents through long tasks
DreamX Team, Alibaba Group · Sept 2026
- Top Universitycs.CL
LLM agents struggle to infer wellbeing from long-term wearable data
Dartmouth College, University of Cambridge · Aug 2026
- Big Techcs.AI
Replayable news timelines test how LLM agents update forecasts
Georgia Institute of Technology, Amazon · Sept 2026
- Top Universitycs.LGcode
New benchmark standardizes evaluation of coding agent harnesses
TokenRhythm Technologies, Infinigence AI · June 2026
- Top Universitycs.CL
Simulated user personas improve agents’ handling of multi-turn workflows
State Key Laboratory of Multimedia Information Processing, Peking University · Sept 2026
- Top Universitycs.AI
New benchmark reveals coordination as distinct bottleneck for LLM agents
University of Edinburgh, University of Oxford · June 2026
- Independentcs.CL
LLM agents do not reliably improve from their own experience
Sept 2026
- Independentcs.CL
Models struggle to improve themselves from vague capability goals
Sept 2026
- Chinese Techcs.AI
Process-level tests expose distinct mathematical agent capabilities in LLMs
Sun Yat-sen University, Tencent Youtu Lab · Aug 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.