The authors introduce OverAct, a benchmark covering 720 episodes across eight privacy-sensitive domains.
Tool-calling agents access more private data than requests require
Across seven models, the authors found systematic over-authorization and reduced privacy-oriented excess by 43 percent using a pre-execution audit.
Chinese Tech
Taolin Zhang · Jiuheng Wan · Hanyu Wang · Tingyuan Hu · Chengyu Wang
Hefei University of Technology · East China Normal University · Alibaba Cloud Computing
Research Digest··2 min read
Zhang et al.
Why this paper
From Alibaba Cloud Computing and 2 others
In one line
LLM tool-calling agents significantly exceed authorized data access, and SelfAudit reduces privacy excess by 43%.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§