The authors propose SCOUT, a two-stage verifier for computer-use agent safety.
Two-stage verifier combines reasoning and tool-use for safer computer agents
SCOUT generates task-specific rubrics through reasoning, then uses tools to inspect environments, outperforming existing verifiers on two benchmarks.
Top University
Jianxing Chen · Xiao Yu · Shipra Agrawal · Zhou Yu
Columbia University
Research Digest··2 min read
The authors introduce SCOUT, a two-stage agentic safety verifier for computer-use agents (CUAs).
Why this paper
From Columbia University · Released code
In one line
SCOUT improves computer-use agent safety by combining reasoning-intensive rubric generation with tool-intensive evidence gathering to detect harmful behaviors.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§