The authors assign an orthogonal projector to each linear layer, or to groups of layers receiving the same activations.
End-to-end subspace learning preserves transformers under aggressive low-rank compression
Learnable Subspace Projections jointly optimize which activation directions to retain, improving accuracy, decoding speed and memory use after factorization.
Top University
Massimo Bini · Anders Christensen · Stephan Alaniz · Judah Goldfeder · Ole Winther · Yann LeCun · +2 more
Helmholtz Munich · Technical University of Munich · MCML · Orbital Industries · LTCI
Research Digest··2 min read
Bini et al.
Why this paper
From Helmholtz Munich and 10 others
In one line
Learnable Subspace Projections (LSP) learns low-rank compression subspaces jointly for better transformer efficiency.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§