The authors built SkillCascade, a multi-agent red-teaming framework that finds workflows involving several skills, modifies those skills to create a coordinated attack, generates test cases, and judges whether the resulting behavior is harmful.
Benign-Looking Agent Skills Can Combine Into Harmful Cross-Skill Attacks
The authors demonstrate attacks whose malicious behavior emerges only when separately modified skills execute together in a shared agent context.
Top University
The Chinese University of Hong Kong, Shenzhen · University at Buffalo, SUNY · University of Oxford
Research Digest··2 min read
Zhu et al.
Why this paper
From The Chinese University of Hong Kong, Shenzhen and 2 others
In one line
Skill cascading attacks distribute a malicious objective across multiple skills so each passes per-skill scans but combined execution causes harm.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§