The authors define deceptive behavior as a model asserting what it does not hold true, a response conditioned on context rather than a fact.
Machines can unlearn deception itself, not just the knowledge it distorts
A contrastive forget set built from the model's own behavior, trained with pressure-aware targets, cuts held-out deceptive responses from over 50% to under 3% while preserving context following.
Academic
Haoran Tang · Rajiv Khanna
Purdue University
Research Digest··3 min read
The authors propose PACT, a machine unlearning objective that removes context-triggered deception from LLM weights rather than deleting the underlying knowledge.
Why this paper
From Purdue University
In one line
PACT unlearns deceptive behaviors in LLMs by training with pressure-aware counterfactual targets, reducing deception to under 3%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§