The authors designed BSD, which treats a frozen LLM as its own teacher.
LLM's hidden beliefs about users can be read and rewritten
Belief Self-Distillation recovers and causally edits a compact user representation, showing refusal depends on inferred user intent and that models share user-state geometry.
Research Lab
Ali Holmov · Yiran Huang · Kirill Bykov · Zeynep Akata
Technical University of Munich · Helmholtz Zentrum München
Research Digest··3 min read
Thread:Hidden-State Steering
Holmov et al.
Why this paper
From Helmholtz Zentrum München and Technical University of Munich · Part of Hidden-State Steering, now 5 papers
In one line
Belief Self-Distillation reveals that LLMs have causally actionable user beliefs, and changing them alters refusal even when the request is fixed.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§