The authors used reinforcement learning to train an attacker LLM to generate prompt injections.
Curriculum training exposes prompt injection flaws in frontier language models
Training an attacker against progressively stronger targets overcame sparse rewards and produced attacks that transferred to previously unseen models.
Academic
Chenlong Yin · Xiaolong Jin · Wei Zou · Yanting Wang · Jinyuan Jia
The Pennsylvania State University · Purdue University
Research Digest··2 min read
Thread:Agent Security & Attacks
Yin et al.
Why this paper
From The Pennsylvania State University and Purdue University · Part of Agent Security & Attacks, now 43 papers
In one line
Curriculum reinforcement learning trains an attacker LLM against progressive targets to attack frontier LLMs where prior methods fail.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§