The authors design a multi-precision quantization scheme called StagQ, building on classical embedded quantization.
Multi-precision LLM weight format keeps a single prefix-readable stream
StagQ stores weights as a 2-bit base with refinement planes, using a side record for outliers, and outperforms baselines on MMLU at 2 and 3 bits.
Chinese Tech
Zhe Wei · Mengqi Guo · Yuan Yuan · Jiunn Bin Lim · Boyi Pan · Michael Bi Mi
Huawei Technologies Ltd.
Research Digest··3 min read
The authors present StagQ, a multi-precision weight format for LLMs that stores weights in a single prefix-readable stream: a 2-bit group-wise affine base followed by configurable 1-bit refinement planes.
Why this paper
From Huawei Technologies Ltd.
In one line
StagQ stores LLM weights in a multi-precision format where any prefix is a valid lower-precision code, improving accuracy at low bits.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§