Pruning experts by replaceability, not magnitude, preserves reasoning in MoE models
The authors introduce RAZOR, a training-free pruning method for mixture-of-experts models that scores experts by how well surviving experts can compensate for their removal. Across four MoE models (GLM-4.7-Flash, Qwen3.6-35B-A3B, DeepSeek-V4-Flash-0731, Hy3) at 25% and 50% pruning budgets, RAZOR achieved the highest macro average over nine reasoning tasks, outperforming the REAP baseline by up to 5.59 points. However, pruned models still exhibited shifts in response diversity, formatting, and termination.