Sensitivity-aware residual-stream pruning improves LLM compression without retraining

The method consistently reduces calibration KL divergence and improves perplexity and downstream task performance across several instruction-tuned model families.

Academic
Chayne Thrash · Kevin Chen · Soheil Kolouri

Vanderbilt University

Research Digest··3 min read
The authors introduce a sensitivity-aware approach to residual-stream pruning for large language models.

Existing residual-stream pruning methods reduce the hidden dimension by minimizing activation reconstruction error, assuming all perturbation directions are equally harmful.

Why this paper

From Vanderbilt University

In one line

Residual-stream pruning that weights activation covariance by output sensitivity preserves LLM output behavior better than activation-only pruning.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.