Language models

Pretraining, architectures, post-training and instruction tuning, reasoning in language models, tokenisation, long context

66 articles · page 1 of 2

Language models

Effective depth shows transformer layers are redundantly correlated

Gahtan et al. introduce effective depth, a scalar that treats a transformer's residual stream as a discrete-time process and measures how representation similarity decays with layer distance. Across 16 decoder-only LLMs, they find that effective depth is almost always below a closed-form reference that assumes maximally diverse orthogonal updates, suggesting that residual updates are strongly correlated.

3 Oct·3 min
Language models

Teacher alignment method closes capacity gap in reasoning distillation without discarding data

The authors propose Teacher Alignment to address the Gap Curse in reasoning distillation, where larger teacher models produce distributions too complex for smaller students to approximate. They develop TeacherGRPO, built on Group Relative Policy Optimization, which adapts the teacher via curriculum selective alignment and importance-adaptive length regularization. Experiments show consistent performance gains over existing baselines across diverse reasoning benchmarks.

3 Oct·3 min
Language models

Language models can suddenly and repeatedly switch between pattern-matching and genuine generalization during pre-training

The authors developed a behavioral test suite to distinguish whether language models rely on shallow patterns or true generalization. Tracking models across pre-training checkpoints, they found that models frequently and abruptly alternate between these modes, often within tens of billions of tokens. This mode-hopping is not an artifact of noise or metric selection and cannot be smoothed by checkpoint averaging.

3 Oct·3 min
Language models

Small models can steer stronger ones through shared reasoning traces

Liang and colleagues introduce Allspark, a framework for transferring reasoning behavior from a reinforcement-learned small model to stronger models through alternating text-based chains of thought. In experiments using Qwen models and ARC-AGI-2, the trained weak teacher improved some within-family and cross-family students, although the benefit depended on the inference configuration and required additional tokens.

3 Oct·3 min
Language models

Geometry-aware LoRA optimization improves convergence across supervised and reinforcement learning

Ding, Zazo and Hensman introduce Rotated Manifold Optimization, or RoM, an optimizer designed around a symmetry in low-rank adaptation: many pairs of LoRA factors represent exactly the same weight update. Across supervised fine-tuning and reinforcement learning experiments, RoM converged faster and reached lower held-out loss than the compared LoRA optimizers, while producing better or comparable downstream performance.

3 Oct·3 min
Language models

Future-aware distillation corrects teacher-student information mismatch in block diffusion language models

The authors introduce d-OPD, a future-aware on-policy distillation method for converting autoregressive LLMs into block diffusion language models. It addresses a fundamental mismatch: the block-diffusion student sees future context within a block, but the standard AR teacher distribution does not. By aligning the teacher distribution with the student's visible state, d-OPD improves average benchmark scores by up to 4.0 points and reduces training time by about 1.35-1.58x on Qwen3 models from 0.6B to 8B.

3 Oct·3 min
Language models

Early answer confidence reveals when language models take reasoning shortcuts

Zhaohan Zhang and colleagues propose ConfLens, a framework that monitors how an LLM's confidence in its final answer evolves during chain-of-thought reasoning. They find that shortcut reasoning consistently shows a pattern of premature confidence, where the model commits to an answer early. Their new metric, the Distributional Answer Commitment Score (DACS), detects this reliably across math and code tasks, improving shortcut detection F1 by over 4.3% over strong baselines.

3 Oct·2 min
Language models

Tabular foundation models refine representations in-context with attention-gated updates

The authors develop in-situ representation refinement, where support labels guide changes to the episode's representations during a forward pass. They propose RefineICL, an attention-gated, FFN-free contextual stack that achieves 0.93836 OVR-AUC and 0.87173 accuracy on AMLB29 and 1644.8 Elo on TabArena, 31.4 above TabPFN-3. A direct intervention confirms that evolving support states are essential: removing one intermediate support update increases final query cross-entropy in all 72 tested episodes.

3 Oct·2 min
Language models

Token-level certainty predicts LLM question difficulty better than response correctness

The authors systematically evaluate whether token-level certainty scores from LLMs predict correctness, distinguishing between identifying questions a model will likely answer correctly and distinguishing correct from incorrect responses to the same question. They find that certainty is generally better at the former, and that the timing of useful information differs: question difficulty appears early, answer correctness appears late. A test-time compute strategy built on these findings raises accuracy from 78.71% to 79.54% while cutting generated-token cost by 82.4%.

2 Oct·3 min
Language models

Value-guided search improves math reasoning during RL training

The authors propose APIVIS, a training-time framework that integrates finite-budget Gumbel search into reinforcement learning with verifiable rewards (RLVR) for mathematical reasoning. By combining direct and searched responses in each rollout group and applying selective supervision to search-improved tokens, they achieve substantial gains over existing search-based methods across benchmarks and model scales.

2 Oct·2 min
Language models

Latent feedback helps Transformers carry information across generation steps

Tirosh, Amos and Geva introduce LIFT, a Transformer architecture that feeds a compact continuous state from each generated position into the processing of the next. Across models from 135 million to 1 billion parameters, the authors report better language modeling, reasoning and procedural-task performance than standard Transformers trained on the same number of tokens, with competitive or better results when training compute is matched.

30 Sept·3 min
Language models

Companion model sharpens LLM confidence without harming task performance

The authors propose concurrent confidence calibration, where confidence is learned alongside capability improvement during reinforcement learning from verifiable rewards (RLVR). Their method, CoCal (Companion Confidence Calibration), trains a separate lightweight model on rollout hidden states and verifier feedback, decoupling confidence estimation from policy optimization. Experiments on Qwen3-8B and Qwen3-14B show that CoCal improves confidence estimation without sacrificing task accuracy, outperforming both RL-based concurrent methods and post-hoc calibration.

30 Sept·3 min
Language models

Training a concept generator improves LLM reasoning over repeated sampling.

The authors propose a method where a small concept generator is trained using reinforcement learning to produce high-level concepts (e.g., hints, strategies) that maximize the downstream success of a frozen, larger answer generator. On hard mathematical reasoning problems, this learned search policy substantially improves pass@k over naive repeated sampling at the same answer generation allocation, transfers to other answer generators, and outperforms concepts from much larger untuned models.

30 Sept·2 min
Language models

Token-level teacher routing beats prompt-level routing in multi-teacher distillation

The authors propose MOPD-Router, a framework for multi-teacher on-policy distillation that replaces prompt-level hard routing with token-level routing across all teachers, requiring no domain labels and no separate routing model. Within this framework they introduce ExpertAlign, a metric that scores each teacher by whether its correction to the student expresses the specialization the teacher acquired during post-training. On unlabeled data, ExpertAlign improves overall scores by 12.3% over Mean aggregation; on domain-labeled data, it beats standard MOPD by 7.8% without using the available domain labels.

30 Sept·2 min
Language models

Structured reasoning framework improves LLM game-playing efficiency and payoff

The authors propose SAGE, a training-free inference-time framework that structures LLM strategic reasoning around three coordinated operations: anchor, adapt, and recalibrate. In repeated imperfect-information games (Leduc Hold'em, Liar's Dice, Goofspiel), SAGE outperforms reasoning-intensive LLM agents, achieving up to a 127.6% payoff improvement in Liar's Dice while reducing input and output token usage by up to 80% and 90% respectively.

30 Sept·3 min