Skill Images Can Steer Agents Into Unauthorized Actions

A benchmark across nine agent configurations found that malicious instructions embedded in visual skill references were more effective than matched text-based attacks.

Chinese Tech
Lingqi Jiang · Jialuo Chen · Jianan Ma · Xinhao Deng · Xiaohu Du · Sibo Yi · +5 more

Zhejiang University · Ant Group · Hangzhou Dianzi University · Tsinghua University

Research Digest··2 min read
Jiang et al.

The authors constructed MMSkillRisk from 28 curated clean skills, producing 36 attack packages and 108 executable cases across five attack objectives.

Why this paper

From Ant Group and 3 others · Released code

In one line

Image-borne attacks in multimodal agent skills induce unauthorized operations in every tested configuration, with a 43.1% success rate.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.