Evaluation & benchmarks

New benchmarks, evaluation methodology, measuring what models can do

97 articles · page 1 of 3

Evaluation & benchmarks

LLMs fail multi-step math despite solving each step correctly

Wang et al. introduce OracleLadder, a diagnostic tool that provides increasing levels of oracle help to large language models on multi-step math problems. By giving models a roadmap of sub-goals and their answers, they isolate failures into five categories. The dominant failure mode across all tested models is a composition gap: models can solve each intermediate step alone but cannot compose them into the correct final answer.

3 Oct·2 min
Evaluation & benchmarks

New benchmark tests automated theorem proving for research-level theoretical computer science

The authors introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery in theoretical computer science. It consists of 398 theorem-level challenges from 138 papers from STOC and COLT 2026. In their evaluations, the best-performing system (GPT-5.6 Sol max with discussion) achieved 23.6% verifier-accepted coverage after 10 rounds, while an agentic planning workflow reached 25.4%.

3 Oct·2 min
Evaluation & benchmarks

New benchmark tests if AI can track objects after they vanish from view

Ma et al. introduce Beyond3D, a VQA benchmark with 9,000 questions over 135 egocentric videos that tests whether VLMs can track an object's location after it is moved and has left the field of view. The best model achieved only 42.2% accuracy, with the largest failures in recovering when an object was last visible, indicating that current models struggle with out-of-sight spatiotemporal reasoning.

3 Oct·2 min
Evaluation & benchmarks

AI agents reproduce only 41% of machine learning papers with full code and weights provided

The authors introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers with predefined reproduction targets and GPU-hour budgets. They test four AI agents once per paper; the best agent reproduces 41% of Run-tier papers (code, data, weights provided), 27% of Retrain-tier (no weights), and 15% of Reimplement-tier (no code). Failed attempts use only 29% of their budget on average, and the most common error is implementing methods without verifying against the paper's reported numbers.

3 Oct·3 min
Evaluation & benchmarks

Video models struggle to connect cinematic techniques with storytelling intent

Xing and colleagues introduce CinematicVQA, a benchmark for reasoning about how cinematographic choices shape what viewers perceive and how a scene functions narratively. Their evaluations reveal a semantic gap: large vision-language models describe visible presentation more reliably than they identify the filming techniques behind it, while targeted fine-tuning improves narrative and multi-step reasoning.

3 Oct·2 min
Evaluation & benchmarks

LLMs struggle to match multidisciplinary tumor board discussions in new benchmark

The authors introduce OpenTumorBoard, a benchmark derived from 12,534 minutes of publicly available tumor board recordings on YouTube, comprising 611 patient cases and 19,157 discussion turns across ten specialist roles. Evaluating 14 general-purpose and medical LLMs, they find that the best models achieve modest scores of 3.43/5 on responding to individual specialist questions and 2.78/5 on generating entire board discussions that align with recorded consensus. Supervised finetuning and reinforcement learning on the benchmark data improve performance on a held-out test set.

3 Oct·3 min
Evaluation & benchmarks

Standard benchmarks overstate the accuracy of reused language model caches

Cestola et al. examine position-independent key-value cache reuse, a technique intended to accelerate retrieval-augmented generation by reusing cached representations of text chunks at different prompt positions. They find that common aggregate accuracy metrics and existing datasets can conceal the accuracy loss caused specifically by reuse, then introduce Boxoffice to generate more demanding evaluations.

3 Oct·3 min
Evaluation & benchmarks

New benchmark measures whether AI systems can explore, not just recall, unfamiliar rules

The authors introduce ExplorationBench, a benchmark that evaluates AI systems' ability to explore unknown environments by forming hypotheses, designing experiments, and iterating on results. Its two executable sandboxes make every answer exactly checkable while ensuring that pretraining recall alone cannot solve the tasks. Evaluating ten AI systems, the authors find that the strongest can acquire and apply these unfamiliar rules, though performance varies across trajectories and continued exploration can stall or reverse earlier gains.

3 Oct·3 min
Evaluation & benchmarks

Clinical language models diagnose well but often mishandle full encounters

Fang et al. introduce a clinician-authored benchmark that evaluates the process of clinical assessment rather than diagnosis from a complete case summary. Across 31 models, diagnostic accuracy reached 90.7%, yet even the strongest systems passed fewer than 30% of tasks when required to question a virtual patient, request appropriate examinations, respect constraints and provide the correct diagnosis.

1 Oct·3 min
Evaluation & benchmarks

Security agents lose evidence support when available telemetry changes

The authors introduce APTInvestBench, a benchmark that tests whether LLM agents can investigate advanced persistent threats when the available security logs change. Across eleven models, agents retrieved sufficient evidence for 44.3% of recoverable attack actions on average, but their final citations adequately supported only 25.0%, revealing substantial losses between finding evidence and reporting it.

1 Oct·2 min
Evaluation & benchmarks

Long context windows do not ensure reliable sustained execution

Willette, Puvvada and Ginsburg introduce Long-Transduction, a diagnostic for whether models can repeatedly read state, transform it and emit correctly aligned outputs over long generations. Their experiments show that advertised context capacity does not imply dependable execution: accuracy fell substantially as context grew from 4K to 128K tokens, even though the underlying operations remained controlled.

1 Oct·2 min
Evaluation & benchmarks

Clinical agents falter when health records cannot support the question

Qiao et al. built an interactive benchmark from MIMIC-IV records to test clinical agents on paired questions that are either supported by the database or rendered unanswerable by a single evidence discrepancy. Across proprietary and large open-weight models, average task success fell from 62.2% on supported questions to 37.9% on noisy counterparts, showing that retrieval and query execution do not ensure evidence-grounded reasoning.

1 Oct·3 min
Evaluation & benchmarks

New benchmark disentangles sycophancy from empathy in multi-turn LLM conversations

The authors introduce FIGS, a benchmark with 500 multi-turn scenarios and an automated judge that separately scores sycophancy (factual yielding) and calibrated validation (appropriate empathy). Evaluating leading LLMs, they find a consistent trade-off: models gradually become sycophantic over sustained conversation, or shift to robotic dismissal, showing that balancing truthfulness with supportiveness remains unsolved.

1 Oct·2 min
Evaluation & benchmarks

Models struggle to predict software behavior after stateful interventions

Xinran Zhang introduces CTE-Bench, an executable benchmark that isolates whether models can predict the downstream effects of modifying a stateful software service. Across four API-hosted models, exact response accuracy on intervention-affected calls reached 54.3% to 61.5% when correct previous responses were supplied, but fell below 34% without that feedback or when models relied on their own predictions.

30 Sept·2 min
Evaluation & benchmarks

Coding agents struggle to turn software proposals into working changes

Peng et al. introduce LoLBench, a benchmark that asks coding agents to implement substantial software enhancements from human-written proposals rather than detailed coding specifications. The results separate two challenges: understanding how a high-level design maps onto an existing repository, and correctly implementing the resulting changes. Even the strongest evaluated agent resolved only 14 percent of tasks.

30 Sept·3 min
Evaluation & benchmarks

Sparse annotation overlap makes LLM judge validation decisions unreliable

Li, Mukherjee and Pal analyze how limited overlap, the proportion of items labeled by multiple raters, affects validation of LLM judges against humans. Across theoretical analysis and experiments with 10 judges, they find that sparse overlap sharply increases erroneous approval, rejection and ranking decisions, while stratified allocation can improve reliability without adding annotations.

30 Sept·2 min
Evaluation & benchmarks

Adaptive testing exposes major security repair gaps in coding agents

The authors developed SecProbe, an evaluation framework that adaptively selects or synthesizes repository-scale vulnerability-repair tasks using Item Response Theory, a statistical method for matching test questions to ability. Across nine frontier models and two agent harnesses, performance remained low, while adaptive testing produced comparable ability estimates with fewer attempted tasks than baseline evaluation strategies.

30 Sept·2 min
Evaluation & benchmarks

LLMs often diagnose without correctly valuing available medical evidence

Feng et al. tested nine language models on MedEVM, where medical evidence arrives sequentially and models must choose whether to wait or diagnose. The authors find that models often submit diagnoses at the wrong evidential moment, remain sensitive to evidence order and misleading details, and gain 12.0 to 51.1 percentage points in accuracy when an evidence-verification harness controls submission.

30 Sept·2 min