Informed attackers can evade static and LLM-based agent skill scanners

Pretext repeatedly rewrites malicious agent skills, achieving up to 97 percent evasion against a fixed scanner and 77 percent against an adapting defense.

Chinese Tech
Tobias Kaisar · Aritra Dhar

Huawei Research Zurich

Research Digest··2 min read
Kaisar and Dhar test whether hybrid skill scanners can withstand an attacker who knows how their checks work.

The authors built a white-box attacker around the design of SkillSpector, a framework that combines static analysis with an LLM judge.

Why this paper

From Huawei Research Zurich

In one line

An informed attacker defeats static and LLM skill scanners by hiding payloads in natural language and splitting instructions across files, succeeding up to 97%.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.