The authors propose a decoding-aware calibration method for training-free weight pruning of large language models.
SparseDecoding aligns pruning calibration with autoregressive decoding activations for faster LLM inference.
The method uses tokens generated by the dense model during decoding to construct calibration matrices, and pairs this with an optimized N:M sparse matrix-vector kernel to achieve up to 1.48x end-to-end speedup on A100 GPUs without sacrificing generation quality.
Chinese Tech
Qitong Wang · Xinwei Niu · Mingluo Su · Shanwei Zhao · Shiai Zhu · Huan Wang
Westlake University · Ant Group
Research Digest··3 min read
Wang et al.
Why this paper
From Ant Group and Westlake University
In one line
Decode-time activation calibration and an N:M SpMV kernel improve pruned LLM generation quality and deliver up to 1.48x decoding speedup on A100 GPUs.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (hardware A100 GPUs)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§