Existing residual-stream pruning methods reduce the hidden dimension by minimizing activation reconstruction error, assuming all perturbation directions are equally harmful.
Sensitivity-aware residual-stream pruning improves LLM compression without retraining
The method consistently reduces calibration KL divergence and improves perplexity and downstream task performance across several instruction-tuned model families.
Academic
Chayne Thrash · Kevin Chen · Soheil Kolouri
Vanderbilt University
Research Digest··3 min read
The authors introduce a sensitivity-aware approach to residual-stream pruning for large language models.
Why this paper
From Vanderbilt University
In one line
Residual-stream pruning that weights activation covariance by output sensitivity preserves LLM output behavior better than activation-only pruning.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§