The authors modified only the decoder of a sparse autoencoder while freezing its encoder and all language-model weights.
Compromised sparse autoencoders can backdoor unchanged language models
Decoder-only modifications caused three frozen language models to insert attacker-selected code, either routinely or in response to a prompt cue.
Academic
Enrico Ahlers · Daniel Passon · Tobias Kiecker · Eik Reichmann · Lars Grunske
Humboldt-Universität zu Berlin
Research Digest··3 min read
Ahlers and colleagues show that sparse autoencoders, auxiliary components used to interpret or steer language models, can carry behavioral backdoors without altering the underlying model.
Why this paper
From Humboldt-Universität zu Berlin
In one line
A malicious SAE decoder can backdoor an unchanged language model, causing unsolicited or prompt-triggered code insertion while conventional SAE quality metrics change only slightly.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§