Lower reconstruction loss can degrade quantized LLM performance; robust refinement fixes it

DRQ refines integer codes to minimize worst-case loss over activation distribution shifts, improving six PTQ methods without inference overhead.

Top University
Yanlong Zhao · Xiaoyuan Cheng · Huihang Liu · Baihua He · Xinyu Zhang · Harrison Bo Hua Zhu · +3 more

University of Science and Technology of China · University College London · Shanghai University of Finance and Economics · AMSS, Chinese Academy of Sciences · University of Copenhagen

Research Digest··2 min read
The authors show that minimizing reconstruction loss on calibration data during weight-only post-training quantization (PTQ) can hurt model performance, even on the very same data.

2-1B-Instruct with 3-bit AWQ using WikiText-2 (WT2) as calibration data.

Why this paper

From AMSS, Chinese Academy of Sciences and 7 others

In one line

Distributionally robust quantization (DRQ) minimizes worst-case reconstruction loss under constrained activation shifts, avoiding the calibration-loss trap.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe