The authors propose H-SVDQuant, a unified low-rank-assisted W4A4 post-training quantization (PTQ) method.
H-SVDQuant improves 4-bit quantization of diffusion transformers with low-rank branches
A unified framework jointly modeling weight and activation quantization error achieves better performance with rank 4 than existing methods with rank 32.
Chinese Tech
Shiwen Wang · Pengxiang Zhao · Xiaoming Yuan
The University of Hong Kong · Huawei Technologies Co., Ltd.
Research Digest··2 min read
Wang et al.
Why this paper
From Huawei Technologies Co., Ltd. and The University of Hong Kong
In one line
H-SVDQuant achieves rank-4 W4A4 quantization that surpasses rank-32 SVDQuant on diffusion transformers and LLMs.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§