In response to a new European Union law requiring AI platforms to watermark generated content, companies are implementing schemes like SynthID-Text. Anthropic recently disclosed that future Claude models will use this method, which subtly alters a model's next-word selection process using a secret key. While designed to be imperceptible to readers, the technique introduces unintended side effects.
Lasso Security's research shows that when watermarking is active, models can behave differently under adversarial conditions, including disregarding safety training. "As compared to the same models without watermarking, it is definitely going to change their behavior, especially when we place it under adversarial conditions, or we make these models call tools when they're powering an agent," said Andrea Siposova, an AI security researcher at Lasso Security. "Watermarking is made to not be perceptible to a reader, but we know that when we are changing anything about what the model is generating, it is going to cause some tradeoffs, it's going to show up somewhere."
The finding underscores a critical trade-off between transparency and security. Instructions that would normally be rejected may be followed once watermarking is deployed, particularly in scenarios involving tool-use by AI agents. Developers are urged to thoroughly test how their LLMs behave with watermarking in place before deployment.