102 papers this week in Evaluation & benchmarks12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Evaluation & benchmarks research

Evaluation & benchmarks

LLMs fail multi-step math despite solving each step correctly

Wang et al. introduce OracleLadder, a diagnostic tool that provides increasing levels of oracle help to large language models on multi-step math problems. By giving models a roadmap of sub-goals and their answers, they isolate failures into five categories. The dominant failure mode across all tested models is a composition gap: models can solve each intermediate step alone but cannot compose them into the correct final answer.

today
Language modelsGOOGL ▲1.6%

Language models can suddenly and repeatedly switch between pattern-matching and genuine generalization during pre-training

The authors developed a behavioral test suite to distinguish whether language models rely on shallow patterns or true generalization. Tracking models across pre-training checkpoints, they found that models frequently and abruptly alternate between these modes, often within tens of billions of tokens. This mode-hopping is not an artifact of noise or metric selection and cannot be smoothed by checkpoint averaging.

today
Evaluation & benchmarks

New benchmark tests automated theorem proving for research-level theoretical computer science

The authors introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery in theoretical computer science. It consists of 398 theorem-level challenges from 138 papers from STOC and COLT 2026. In their evaluations, the best-performing system (GPT-5.6 Sol max with discussion) achieved 23.6% verifier-accepted coverage after 10 rounds, while an agentic planning workflow reached 25.4%.

today

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses