The authors derived a theoretical optimal linear compression for influence functions: top-k PCA coordinates of half-whitened training gradients.
Compressed gradient representations preserve influence estimates in large language models
Eigenbasis-corrected one-bit gradient projection (EOGP) stores gradients with 16x less storage while matching or exceeding baselines on GPT-2 and OLMo models.
Top University
Jaeseung Heo · J Rosser · Dongwoo Kim
POSTECH · University of Oxford
Research Digest··2 min read
The authors propose EOGP, a method to compress training gradients for reusable influence function computation in LLMs.
Why this paper
From University of Oxford and POSTECH
In one line
Eigenbasis-corrected one-bit gradient projection compresses gradients for influence functions, using one sixteenth the storage of baselines while preserving accuracy.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§