The authors tested whether decoding-time attacks, which alter a model’s next-token choices during generation, can work through text-only interfaces.
Sampled text can steer black-box models past safety safeguards
BlindBias reconstructs and selectively modifies decoding distributions using repeated text samples, without model weights or numerical token probabilities.
Big Tech
Jesson Wang · Shawn Li · Wei Yang · Franck Dernoncourt · Ryan A. Rossi · Charith Peris · +1 more
University of Southern California · Adobe · Amazon
Research Digest··2 min read
Wang et al.
Why this paper
From Amazon and 2 others
In one line
Text-only LLM interfaces remain vulnerable to decoding-time jailbreaks when they allow repeated sampling and continuation from supplied assistant prefixes.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§