The authors normalize each weight-matrix row, group weights into blocks of eight and quantize them onto the E8 lattice, a structured set of points in eight-dimensional space.
EntroPack compresses neural weights at finely adjustable storage rates
The method combines lattice quantization, sampled rate estimation and tiled GPU decoding to meet target bitrates without calibration or fine-tuning.
Independent
Hong Zhang · Zhongjie Duan · Yingda Chen
Research Digest··2 min read
Zhang, Duan and Chen present a weight compressor that can target rates between the coarse options offered by fixed-width formats.
Why this paper
Independent · Released code
In one line
EntroPack entropy-codes row-normalized E8 lattice quantized weights to hit arbitrary bitrates with lower error than fixed-width formats and fast GPU reconstruction.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§