The authors studied eight open-weight LLMs (four from the Llama 3 family, two from Mistral, one from Gemma, and one from Qwen).
Tool use shifts how LLMs compute refusal, not just their outputs
Rodríguez et al. show that harmful requests are still internally perceived as harmful, but the mechanism converting that perception into a refusal is weaker and more brittle when models interact via tools.
Top University
Abel Rodríguez · Giuseppe Garofalo · Lieven Desmet · Vera Rimmer
KU Leuven
Research Digest··3 min read
The authors systematically compare how large language models (LLMs) process and refuse harmful requests in conversational versus tool-mediated settings.
Why this paper
From KU Leuven
In one line
Tool mediation raises the refusal threshold and makes refusal more brittle in large language models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§