The authors evaluated tool-using and no-tool configurations on MM-SafetyBench, HoliSafe and VLSBench, three benchmarks pairing images with potentially harmful requests.
Visual language models refuse harmful requests less often when using tools
Across three safety benchmarks, adding agentic visual tools consistently increased refusal failures in open-weight and proprietary models.
Big Tech
Rikiya Takehi · Ryo Hachiuma · Shaona Ghosh · Dan Zhao · Yu-Chiang Frank Wang · Yusuke Hirota
MIT · NVIDIA
Research Digest··3 min read
Takehi et al.
Why this paper
From NVIDIA and MIT · Part of Agent Security & Attacks, now 60 papers
In one line
Agentic MLLMs fail to refuse harmful requests when using tools, with refusal failure increasing by up to 68.7%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§