End-to-end subspace learning preserves transformers under aggressive low-rank compression

Learnable Subspace Projections jointly optimize which activation directions to retain, improving accuracy, decoding speed and memory use after factorization.

Top University
Massimo Bini · Anders Christensen · Stephan Alaniz · Judah Goldfeder · Ole Winther · Yann LeCun · +2 more

Helmholtz Munich · Technical University of Munich · MCML · Orbital Industries · LTCI

Research Digest··2 min read
Bini et al.

The authors assign an orthogonal projector to each linear layer, or to groups of layers receiving the same activations.

Why this paper

From Helmholtz Munich and 10 others

In one line

Learnable Subspace Projections (LSP) learns low-rank compression subspaces jointly for better transformer efficiency.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.