The authors define Target Reference Advantage (TRA) as the excess token log-likelihood assigned by a model trained on a record relative to a model trained with the same procedure on the complementary half of the corpus.
Tokenwise reference penalties curb privacy memorization during language-model fine-tuning
TRAP suppresses unusually high probabilities for record-specific tokens without requiring sensitive spans to be identified beforehand.
Big Tech
Muhammed Ustaomeroglu · Ziyue Xu · Hanshen Xiao · Peter Cnudde · Guannan Qu · Holger R. Roth
Carnegie Mellon University · NVIDIA · Purdue University
Research Digest··2 min read
Ustaomeroglu et al.
Why this paper
From NVIDIA and 2 others
In one line
TRAP penalty reduces memorization of sensitive data in fine-tuned language models with little utility loss.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§