Tetra encodes each block of 24 weights in 48 bits using the structure of the Golay error-correcting code underlying the Leech lattice.
Tetra serves 2.7-bit language models without expanding lattice codes
A fused GPU kernel decodes compact Leech-lattice weights during multiplication, reducing memory traffic while retaining usable accuracy on Qwen3 models.
Independent
Pier-Jean Malandrino
Research Digest··3 min read
Malandrino presents Tetra, a representation and serving kernel that makes two-bit Leech-lattice quantization practical without first expanding its enormous codebook into a wider GPU format.
Why this paper
Independent
In one line
Tetra serves Leech-lattice-quantized Qwen3 models at about 2.7 bits per parameter by decoding compact codes inside matrix-vector multiplication.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 4B)
- ✓Limitations stated by the authors (3 noted)
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§