SparseDecoding aligns pruning calibration with autoregressive decoding activations for faster LLM inference.

The method uses tokens generated by the dense model during decoding to construct calibration matrices, and pairs this with an optimized N:M sparse matrix-vector kernel to achieve up to 1.48x end-to-end speedup on A100 GPUs without sacrificing generation quality.

Chinese Tech
Qitong Wang · Xinwei Niu · Mingluo Su · Shanwei Zhao · Shiai Zhu · Huan Wang

Westlake University · Ant Group

Research Digest··3 min read
Wang et al.

The authors propose a decoding-aware calibration method for training-free weight pruning of large language models.

Why this paper

From Ant Group and Westlake University

In one line

Decode-time activation calibration and an N:M SpMV kernel improve pruned LLM generation quality and deliver up to 1.48x decoding speedup on A100 GPUs.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (hardware A100 GPUs)
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe