Compromised sparse autoencoders can backdoor unchanged language models

Decoder-only modifications caused three frozen language models to insert attacker-selected code, either routinely or in response to a prompt cue.

Academic
Enrico Ahlers · Daniel Passon · Tobias Kiecker · Eik Reichmann · Lars Grunske

Humboldt-Universität zu Berlin

Research Digest··3 min read
Ahlers and colleagues show that sparse autoencoders, auxiliary components used to interpret or steer language models, can carry behavioral backdoors without altering the underlying model.

The authors modified only the decoder of a sparse autoencoder while freezing its encoder and all language-model weights.

Why this paper

From Humboldt-Universität zu Berlin

In one line

A malicious SAE decoder can backdoor an unchanged language model, causing unsolicited or prompt-triggered code insertion while conventional SAE quality metrics change only slightly.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.