Language models

Pretraining, architectures, post-training and instruction tuning, reasoning in language models, tokenisation, long context

66 articles · page 2 of 2

Language models

Hybrid backbones can efficiently adapt into diffusion language models

The authors adapted pretrained Qwen3.5 models at four parameter scales into diffusion language models, which generate by iteratively filling masked positions rather than strictly proceeding left to right. Despite the backbone’s structurally causal recurrent layers, the resulting dQwen3.5 models retained any-order and parallel decoding behavior while adapting more efficiently than a full-attention baseline.

20 Sept·2 min
Language models

Sparse mixture-of-experts models overfit repeated training data faster

Jha et al. trained dense Transformers and mixture-of-experts models under varying data-repetition rates, domain mixtures, expert counts, and expert granularities. They find that MoE models lose performance more rapidly than dense models as examples are reused, with susceptibility tracking total parameter count rather than the number of parameters activated per token. Strong masking-based regularization preserves an MoE advantage beyond 64 repetitions, but does not match training on unique data.

13 Sept·2 min
Language models

Hierarchical memory trees improve long-context reasoning without model training

Zhang et al. introduce ConvMem, a training-free framework that uses prompted language models as query-dependent summarization kernels over segments of long documents. On two synthetic long-context, multi-hop question-answering benchmarks, the authors report better performance than training-free baselines and stronger out-of-distribution behavior than reinforcement-learning-trained memory methods.

11 Sept·2 min
Language models

Distillation Before Reinforcement Learning Improves Reasoning Model Post-Training

Li and colleagues compare ways to combine on-policy distillation (OPD), which supplies dense teacher feedback at the token level, with reinforcement learning from verifiable rewards (RLVR), which uses sparse outcome-based signals. Across logic and mathematics benchmarks, they find that completing OPD before switching to RLVR works consistently better than pure OPD, pure RLVR, or joint optimization.

5 Sept·2 min
Language models

Graph Machine architecture uses dynamic sparse routing to handle large state efficiently

The authors introduce the Graph Machine (GM), an architecture that maintains an O(n)-sized state and accesses it via sparse, dynamic routing using edges—pointer-like objects updated by a referral mechanism. They replace 75% of the dense Transformer layers in Qwen3-0.6B with GM sparse layers and pretrain from scratch on 15.7B tokens. With only 2 of 4,096 tokens retrieved per KV head per sparse layer, loss degrades only slightly; with 4 tokens, the best model marginally improves loss over the dense baseline.

4 Sept·2 min
Language models

Self-evolving loop synthesizes high-quality multimodal training data

The authors present VISA, an agentic framework that reformulates multimodal instruction synthesis as a self-evolving loop. VISA iteratively generates diverse and challenging instruction-following samples, using verifier signals and target-model failure profiles to guide subsequent rounds. The method consistently improves multimodal instruction following on MM-IFEval while maintaining general multimodal capability across seven benchmarks.

28 Aug·2 min
Language models

Sliding reasoning windows make long test-time scaling 3x faster

Muennighoff et al. observe that during long reasoning traces, most intermediate tokens lose importance as the model continues. They propose Prefix Sliding, which keeps only the instruction prefix and a recent window of tokens, discarding the rest to cap memory usage. Applied without training to existing models, it achieves up to 3x speedup with no performance loss; training with it enables reasoning traces beyond 100,000 tokens.

28 Aug·2 min
Language models

A 2.8-trillion-parameter open MoE model approaches frontier performance.

Kimi Team introduces Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104B activated parameters, native vision, and a 1M-token context window. The authors demonstrate that innovations in attention (Kimi Delta Attention, Attention Residuals) and routing (Stable LatentMoE) yield roughly 2.5× scaling efficiency improvement over Kimi K2, enabling frontier-level performance across long-horizon coding, agentic, reasoning, and vision tasks.

28 July·2 min
Language models

Probabilistic scoring unlocks verification as a new scaling axis for LLMs

The authors introduce LLM-as-a-Verifier, a general-purpose verification framework that extracts continuous scores from LLMs by taking expectations over scoring token logits. This approach achieves state-of-the-art results on four benchmarks—Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%)—and provides fine-grained feedback that can be used for reinforcement learning and progress monitoring.

7 July·3 min