Alternative tokenizations can bypass targeted LLM editing and unlearning
Baser et al. test whether knowledge edits and unlearning remain effective when an adversary changes how an input is divided into tokens, while preserving the underlying text. Across five language models, six datasets and six modification techniques, alternative tokenizations frequently restored the model’s pre-modification response.