LLMs fail multi-step math despite solving each step correctly
Wang et al. introduce OracleLadder, a diagnostic tool that provides increasing levels of oracle help to large language models on multi-step math problems. By giving models a roadmap of sub-goals and their answers, they isolate failures into five categories. The dominant failure mode across all tested models is a composition gap: models can solve each intermediate step alone but cannot compose them into the correct final answer.
3 Oct 2026