PrivDrift tests active-context re-disclosure, where a secret supplied by a user remains in the current conversation and is later requested after unrelated, content-dense turns.
Topic shifts do not reliably stop LLMs revealing user secrets
PrivDrift found substantial secret re-disclosure across three long-context models, even after conversations moved through unrelated topics.
Academic
West Virginia University
Research Digest··2 min read
The author built 1,000 controlled multi-turn dialogues containing seeded secrets, unrelated topic drift and later extraction attempts with varying persuasive styles.
Why this paper
From West Virginia University
In one line
User secrets in LLM conversations remain recoverable after topic drift, with leakage rates between 38.7% and 54.6%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§