Tetra serves 2.7-bit language models without expanding lattice codes

A fused GPU kernel decodes compact Leech-lattice weights during multiplication, reducing memory traffic while retaining usable accuracy on Qwen3 models.

Independent
Pier-Jean Malandrino
Research Digest··3 min read
Malandrino presents Tetra, a representation and serving kernel that makes two-bit Leech-lattice quantization practical without first expanding its enormous codebook into a wider GPU format.

Tetra encodes each block of 24 weights in 48 bits using the structure of the Golay error-correcting code underlying the Leech lattice.

Why this paper

Independent

In one line

Tetra serves Leech-lattice-quantized Qwen3 models at about 2.7 bits per parameter by decoding compact codes inside matrix-vector multiplication.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ✓Compute or model size stated (params 4B)
  • ✓Limitations stated by the authors (3 noted)
  • ✓Reports numbers on named benchmarks (3 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.