CanonQ first maps different tensor sources into a shared representation.
Canonical training improves language models under joint low-bit compression
CanonQ uses fixed transforms and reusable codebooks, then jointly trains models to tolerate quantized weights, activations and attention caches.
Big Tech
Kai Yi · Tarek Elgamal · Sruthikesh Surineni · Vignesh Vivekraja · Soumyadeep Ghosh · Steven Li
Meta AI
Research Digest··3 min read
Yi et al.
Why this paper
From Meta AI
In one line
CanonQ unifies weight, activation, and KV cache quantization using frozen reference codebooks and joint training.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 1B/3B/8B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (4 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§