134 papers this week in Language models12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Language models research

Evaluation & benchmarks

LLMs fail multi-step math despite solving each step correctly

Wang et al. introduce OracleLadder, a diagnostic tool that provides increasing levels of oracle help to large language models on multi-step math problems. By giving models a roadmap of sub-goals and their answers, they isolate failures into five categories. The dominant failure mode across all tested models is a composition gap: models can solve each intermediate step alone but cannot compose them into the correct final answer.

today
Efficiency & systems

Pruning experts by replaceability, not magnitude, preserves reasoning in MoE models

The authors introduce RAZOR, a training-free pruning method for mixture-of-experts models that scores experts by how well surviving experts can compensate for their removal. Across four MoE models (GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, Hy3) at 25% and 50% pruning budgets, RAZOR achieved the highest macro average over nine reasoning tasks, outperforming the REAP baseline by up to 5.59 points. However, pruned models still exhibited shifts in response diversity, formatting, and termination.

today
Efficiency & systems

Training-free inference framework trims redundant computation and memory in looped Transformers

FlashLoop is a training-free inference framework that reduces cross-loop redundancy in looped Transformers through token-sparse updates, sparse attention, and KV-residual quantization. The authors show that as loops progress, state changes concentrate on few tokens, attention-output differences concentrate on few key columns, and KV residuals become quantization-friendly. Across several looped models, FlashLoop achieves lossless accuracy with up to 1.64x speedup and 6x KV-cache memory reduction.

today

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses