The authors ran a low-cost surrogate model alongside a black-box agent.
Small surrogate models can flag errors in black-box LLM agents
A training-free open-weight sidecar ranked tool-call and coding errors more reliably than agents’ self-reported confidence or repeated sampling.
Big Tech
Yikai Zhao · Saurabh Pandey · Pradeep Kumar Misra
Amazon
Research Digest··3 min read
Thread:Agent Security & Attacks
Zhao, Pandey and Misra audit black-box LLM agents by having a smaller open-weight model score each proposed action from its token log-probabilities.
Why this paper
From Amazon · Part of Agent Security & Attacks, now 60 papers
In one line
A small surrogate model's log-probabilities audit black-box LLM agent actions better than the agent's own confidence or self-consistency.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§