EntroPack compresses neural weights at finely adjustable storage rates

The method combines lattice quantization, sampled rate estimation and tiled GPU decoding to meet target bitrates without calibration or fine-tuning.

Independent
Hong Zhang · Zhongjie Duan · Yingda Chen
Research Digest··2 min read
Zhang, Duan and Chen present a weight compressor that can target rates between the coarse options offered by fixed-width formats.

The authors normalize each weight-matrix row, group weights into blocks of eight and quantize them onto the E8 lattice, a structured set of points in eight-dimensional space.

Why this paper

Independent · Released code

In one line

EntroPack entropy-codes row-normalized E8 lattice quantized weights to hit arbitrary bitrates with lower error than fixed-width formats and fast GPU reconstruction.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.