Efficiency & systems

Quantisation, inference serving, kernels, distributed training, compression, efficient architectures

34 articles

Efficiency & systems

Pruning experts by replaceability, not magnitude, preserves reasoning in MoE models

The authors introduce RAZOR, a training-free pruning method for mixture-of-experts models that scores experts by how well surviving experts can compensate for their removal. Across four MoE models (GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, Hy3) at 25% and 50% pruning budgets, RAZOR achieved the highest macro average over nine reasoning tasks, outperforming the REAP baseline by up to 5.59 points. However, pruned models still exhibited shifts in response diversity, formatting, and termination.

3 Oct 2026
Efficiency & systems

Training-free inference framework trims redundant computation and memory in looped Transformers

FlashLoop is a training-free inference framework that reduces cross-loop redundancy in looped Transformers through token-sparse updates, sparse attention, and KV-residual quantization. The authors show that as loops progress, state changes concentrate on few tokens, attention-output differences concentrate on few key columns, and KV residuals become quantization-friendly. Across several looped models, FlashLoop achieves lossless accuracy with up to 1.64x speedup and 6x KV-cache memory reduction.

3 Oct 2026
Efficiency & systems

Masked co-adaptive fine-tuning cuts expert fetches by over 23% for MoE models.

Wu et al. propose MaskCoFT, a fine-tuning method that uses learnable binary masks to restrict each layer's Top-K routing to a subset of experts while allowing those experts to adapt to the redirected tokens. For Mixtral-8x7B and DeepSeek-V2-Lite, the method reduces expert fetches per token by 23.7% and 10.1%, respectively, and lowers time per output token by up to 16.4% and 5.5% in a real offloading system. Average accuracy over nine benchmarks remains slightly above the base model.

3 Oct 2026
Efficiency & systems

Compressed LLM memories can preserve skills beyond their source documents

Laguna and colleagues test whether compressed key-value caches remain safe to reuse when a user asks questions unrelated to the cached document. They find that standard Cartridges preserve document-specific information well but can contaminate unrelated answers, weaken general knowledge and disrupt instruction following. Two simple modifications largely preserve these off-context capabilities with little reported loss of document utility.

3 Oct 2026
Efficiency & systems

Training-free feature folding cuts memory and improves accuracy for wide-table tabular models.

Zhou et al. introduce Support-Compiled Feature Folding (SCFF), a training-free inference method for tabular foundation models that avoids full-width feature mixing by organizing columns into a core and a tail. On an 18-dataset wide-table benchmark, SCFF improved accuracy and negative log-likelihood across all six evaluated frozen backbones, while cutting median peak GPU memory by 2.09-2.36x. The method also used saved memory to retain more evidence, yielding accuracy gains of 4.06 and 3.72 points over widest feasible single-leaf baselines.

3 Oct 2026
Efficiency & systems

Training-free MoE expert pruning via consensus residuals preserves reasoning performance

The authors introduce Razor, a training-free pruning method for mixture-of-experts LLMs that selects experts to remove by aggregating consensus residuals—exact deviations of expert outputs from the weighted mixture. On GLM-4.7-Flash and Qwen3.6-35B-A3B at 25% and 50% expert removal, Razor achieves the highest macro average on nine reasoning tasks, gaining 2.12 to 5.59 points over the REAP baseline while also reducing reverse KL divergence.

3 Oct 2026
Efficiency & systems

Looped Transformer inference sped up by exploiting sparse cross-loop redundancy

The authors demonstrate that much of the computation in looped Transformers is redundant: as recurrence progresses, changes become concentrated on a small subset of tokens, attention-output differences are dominated by sparse key columns, and key-value (KV) residuals between adjacent loops become amenable to low-bit quantization. They introduce FlashLoop, a training-free inference framework that uses token-sparse updates, loop-aware sparse attention, and KV residual quantization, achieving up to 1.64× end-to-end speedup and up to 6× KV-cache memory reduction without loss of accuracy.

3 Oct 2026
Efficiency & systems

Elastic phase sharing improves efficiency in agentic LLM serving

Xu et al. address a mismatch in disaggregated LLM serving, where separate pools handle prompt processing, called prefill, and token generation, called decode, even though demand for each phase changes rapidly. Across public and internal traffic traces, Crossflow increased token throughput by 16.2 to 17.4% on geometric mean over static partitioning, while reducing average time to first token at every evaluated load.

30 Sept 2026
Efficiency & systems

Adaptive prefill-decode execution and elastic GPU reuse boost agentic RL throughput

The authors introduce PEARL, a system that coordinates external elastic GPU resources, temporary reuse of idle training GPUs, and adaptive prefill-decode (PD) configuration for asynchronous agentic reinforcement learning. In experiments with Qwen3-8B and Qwen3-30B-A3B models, PEARL achieves 2.17 to 2.79 times the throughput of a fixed-resource baseline (ROLL) and up to 26.9% and 36.3% improvement over RLBoost+ for the respective models.

30 Sept 2026
Efficiency & systems

Dispatch-aware agents safely optimize compiler-generated GPU kernels within full models

Poddar et al. present a five-agent optimization system that respects compiler dispatch decisions, leaving cuBLAS and cuDNN calls intact while targeting generated Triton kernels. On 250 KernelBench problems, KernelOPT reports geometric-mean speedups over torch.compile ranging from 1.07× to 1.40×, with a verification cascade that retains the baseline whenever no candidate passes.

27 Sept 2026
Efficiency & systems

Engineering agents need evidence-bound authorization before their outputs trigger action

Koch proposes a cross-domain assurance architecture for deciding when agent-produced software, firmware, or PCB artifacts may advance to consequential actions. Rather than treating tests, attestations, or human approval as sufficient in isolation, the model requires evidence to remain artifact-specific, policy-bound, current, independent where risk demands it, and valid at the moment of execution.

25 Sept 2026
Efficiency & systems

Selective attention calls, guided by model's own state, speed long-context inference

The authors introduce On-Demand Attention (ODA), a decoding method that uses a trained recall head to decide when to invoke global attention based on the model's decoding state. ODA trains only the recall head, leaving the pretrained model untouched, and achieves substantial speedups in long-context inference by selectively reading the full history. Experiments show ODA recovers near-full-attention performance while using far fewer global attention calls.

21 Sept 2026