The authors study how context tokens in a transformer-based language model can be absorbed into model weights as a multiplicative operator.
Safety instruction influence controlled via eigenvalue modulation in transformers
Authors show that a safety instruction's spectral operator eigenvalue acts as a continuous dial, enabling Pareto improvements in attack success versus over-refusal tradeoffs.
Big Tech
Benoit Dherin · Michael Munn · Xavier Gonzalvo · Adrian Goldwaser · Blaz Bratanic · Ananth Balashankar · +11 more
Google Research · Cambridge University · Google DeepMind · Google · Tel Aviv University
Research Digest··2 min read
The authors derive a Contrastive Safety Loss with a suppression weight that modulates the dominant eigenvalue of a safety instruction's operator in transformer models.
Why this paper
From Google Research and 4 others
In one line
Safety instructions in transformer LMs act as a multiplicative operator whose dominant eigenvalue controls instruction strength, enabling a contrastive loss to tune the attack/over-refusal tradeoff.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§