Multi-precision LLM weight format keeps a single prefix-readable stream

StagQ stores weights as a 2-bit base with refinement planes, using a side record for outliers, and outperforms baselines on MMLU at 2 and 3 bits.

Chinese Tech
Zhe Wei · Mengqi Guo · Yuan Yuan · Jiunn Bin Lim · Boyi Pan · Michael Bi Mi

Huawei Technologies Ltd.

Research Digest··3 min read
The authors present StagQ, a multi-precision weight format for LLMs that stores weights in a single prefix-readable stream: a 2-bit group-wise affine base followed by configurable 1-bit refinement planes.

The authors design a multi-precision quantization scheme called StagQ, building on classical embedded quantization.

Why this paper

From Huawei Technologies Ltd.

In one line

StagQ stores LLM weights in a multi-precision format where any prefix is a valid lower-precision code, improving accuracy at low bits.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.