56 papers this week in Efficiency & systems12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Efficiency & systems research

Latest Paper· Generative media

Adaptive reference attention reuse speeds up conditioned diffusion transformers

The authors propose RefAdapt-DiT, a training-free framework that adaptively controls joint attention computation between reference images and target generation in diffusion transformers. By monitoring reference drift and target-to-reference attention, it selectively reuses stale reference states, achieving speedups of up to 3.54× on image editing with minimal quality loss.

Jian Tang, Jiawei Fan, Qiannan Zhou +3
Efficiency & systems

Pruning experts by replaceability, not magnitude, preserves reasoning in MoE models

The authors introduce RAZOR, a training-free pruning method for mixture-of-experts models that scores experts by how well surviving experts can compensate for their removal. Across four MoE models (GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, Hy3) at 25% and 50% pruning budgets, RAZOR achieved the highest macro average over nine reasoning tasks, outperforming the REAP baseline by up to 5.59 points. However, pruned models still exhibited shifts in response diversity, formatting, and termination.

today
Efficiency & systems

Training-free inference framework trims redundant computation and memory in looped Transformers

FlashLoop is a training-free inference framework that reduces cross-loop redundancy in looped Transformers through token-sparse updates, sparse attention, and KV-residual quantization. The authors show that as loops progress, state changes concentrate on few tokens, attention-output differences concentrate on few key columns, and KV residuals become quantization-friendly. Across several looped models, FlashLoop achieves lossless accuracy with up to 1.64x speedup and 6x KV-cache memory reduction.

today
Efficiency & systems

Masked co-adaptive fine-tuning cuts expert fetches by over 23% for MoE models.

Wu et al. propose MaskCoFT, a fine-tuning method that uses learnable binary masks to restrict each layer's Top-K routing to a subset of experts while allowing those experts to adapt to the redirected tokens. For Mixtral-8x7B and DeepSeek-V2-Lite, the method reduces expert fetches per token by 23.7% and 10.1%, respectively, and lowers time per output token by up to 16.4% and 5.5% in a real offloading system. Average accuracy over nine benchmarks remains slightly above the base model.

today

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses