The authors first measured how supervised fine-tuning and preference-based defenses alter a model’s output distribution.
Self-distillation hardens tool-using agents without sacrificing benign task performance
RAISED trains agents to ignore injected instructions while preserving legitimate guidance contained in tool outputs.
Big Tech
Mohamed Dhouib · Clement Elliker · Alexi Canesse · Maël Jenny · Lucas-Andrei Thil · Mahammed El-Sharkawy · +2 more
LIX · École polytechnique · Institut Polytechnique de Paris · CNRS · Google DeepMind
Research Digest··2 min read
Dhouib et al.
Why this paper
From Google DeepMind and 6 others · Released code
In one line
RAISED uses self-distillation to defend LLM agents from prompt injection without losing utility.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§