The authors evaluated eight open-weight language-model systems using 264 seed prompts spanning 11 hazard categories.
Jailbreak prompts measurably weakened safeguards across eight open-weight models
MLCommons paired ordinary and adversarial prompts to quantify how much single-turn jailbreaks increased unsafe responses.
Big Tech
Carsten Maple (Victor) · Cagatay Yucel (Victor) · Isaac Holeman (Victor) · Chris Knotz (Victor) · Peter Mattson (Victor) · James Goel (Victor) · +32 more
Research Digest··2 min read
The authors present an end-to-end benchmark for testing whether text-based jailbreak prompts bypass large language model safeguards.
Why this paper
From Google DeepMind and 15 others
In one line
Jailbreaks raise unsafe-response rates from 11.08% to 18.65% across eight open-weight LLMs.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§