The authors built a white-box attacker around the design of SkillSpector, a framework that combines static analysis with an LLM judge.
Informed attackers can evade static and LLM-based agent skill scanners
Pretext repeatedly rewrites malicious agent skills, achieving up to 97 percent evasion against a fixed scanner and 77 percent against an adapting defense.
Chinese Tech
Tobias Kaisar · Aritra Dhar
Huawei Research Zurich
Research Digest··2 min read
Kaisar and Dhar test whether hybrid skill scanners can withstand an attacker who knows how their checks work.
Why this paper
From Huawei Research Zurich
In one line
An informed attacker defeats static and LLM skill scanners by hiding payloads in natural language and splitting instructions across files, succeeding up to 97%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§