The authors built dattri-LLM around two exact gradient representations (factorized and materialized) and a FLOP-aware cost model that selects the cheaper route per layer and operation, sharing kernels across attribution methods.
A library that makes gradient-based training data attribution practical for large language models
dattri-LLM achieves 3.2x speedup over existing tools and scales to 110B-parameter models without modifying training pipelines
Big Tech
Shixuan Liu · Tongli Zhou · Junwei Deng · Pingbang Hu · Jiaqi W. Ma
University of Illinois at Urbana-Champaign · Google
Research Digest··2 min read
The authors introduce dattri-LLM, a library for training data attribution (TDA) at LLM scale.
Why this paper
From Google and University of Illinois at Urbana-Champaign
In one line
A library called dattri-LLM speeds gradient-based training data attribution by 3.2x and scales to 110B-parameter models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§