The authors propose DP-MixMin, an end-to-end pipeline that privately selects the weighting of several public pretraining datasets for a given sensitive downstream task.
Privately learning public data mixtures improves differentially private training
The authors introduce DP-MixMin, a pipeline that privately selects a weighted pretraining mixture, improving chest X-ray AUC by up to 0.037 and reducing language model perplexity by 16%.
Top University
Yufei Chen · Tejumade Afonja · Anvith Thudi · Nicolas Papernot
University of Toronto · Vector Institute · CISPA Helmholtz Center for Information Security
Research Digest··3 min read
Chen, Afonja, Thudi, and Papernot present DP-MixMin, a privacy-preserving method that learns the best mixture of public datasets for pretraining before differentially private fine-tuning on sensitive data.
Why this paper
From CISPA Helmholtz Center for Information Security and 2 others
In one line
A pipeline that privately learns the mixture of public datasets for pretraining improves private downstream task performance.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§