The authors divide an agent attack into two Risk-Realization Factors.
Splitting malicious actions from their rationale defeats skill-by-skill audits
The authors show that coordinated agent skills can fabricate context upstream, making a downstream malicious operation appear justified.
Chinese Tech
Wenxin Wu · Lingyong Yan · Lei Sha · Shuaiqiang Wang · Jiashu Zhao
Beihang University · Baidu Inc · Wilfrid Laurier University
Research Digest··3 min read
Wu and colleagues introduce CoordPoison, an automated framework for poisoning reusable LLM-agent skills by separating a malicious operation from the contextual rationale that prompts its execution.
Why this paper
From Baidu Inc and 2 others · Released code
In one line
Separating a poisoned agent skill's malicious operation from its fabricated justification enables effective attacks that isolated skill audits can miss.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§