The authors studied trainable links that map one language model agent’s internal representations into another agent’s input space.
Learned latent links can undermine safety in multi-agent systems
The authors show that modifying representation-level communication links can increase harmful compliance without changing the safety-aligned agents themselves.
Research Lab
Muhammad Huzaifa · Sina Mavali · Thorsten Eisenhofer
CISPA Helmholtz Center for Information Security
Research Digest··3 min read
Huzaifa, Mavali and Eisenhofer test whether latent communication, where agents exchange internal representations rather than text, creates a distinct safety risk.
Why this paper
From CISPA Helmholtz Center for Information Security · Part of Agent Security & Attacks, now 43 papers
In one line
Training latent communication links can undermine otherwise safety-aligned agents, while reward-guided link optimization can both amplify and repair harmful compliance.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§