Tool use shifts how LLMs compute refusal, not just their outputs

Rodríguez et al. show that harmful requests are still internally perceived as harmful, but the mechanism converting that perception into a refusal is weaker and more brittle when models interact via tools.

Top University
Abel Rodríguez · Giuseppe Garofalo · Lieven Desmet · Vera Rimmer

KU Leuven

Research Digest··3 min read
The authors systematically compare how large language models (LLMs) process and refuse harmful requests in conversational versus tool-mediated settings.

The authors studied eight open-weight LLMs (four from the Llama 3 family, two from Mistral, one from Gemma, and one from Qwen).

Why this paper

From KU Leuven

In one line

Tool mediation raises the refusal threshold and makes refusal more brittle in large language models.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.