The authors created PrivDrift, a benchmark of 1,000 controlled dialogues containing seeded user secrets, content-dense topic-drift turns, and standardized prompts designed to extract the earlier information.
Topic Changes Do Not Reliably Stop LLMs From Revealing User Secrets
PrivDrift tests whether three long-context language models disclose seeded secrets after unrelated dialogue and persuasion-based extraction attempts.
Industry
Luciano Maldonado
West Virginia University
Research Digest··2 min read
Thread:Agent Security & Attacks
The authors built 1,000 controlled, multi-turn conversations in which users disclose secrets before the discussion shifts to unrelated, content-heavy topics.
Why this paper
From West Virginia University · Part of Agent Security & Attacks, now 26 papers
In one line
User-disclosed secrets remain recoverable from active LLM conversations even after topic drift.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§