SecProbe observes an agent's successes and failures, estimates its ability using Item Response Theory, then chooses the next informative task from an existing bank or generates a new task at a targeted difficulty.
Adaptive testing exposes major security repair gaps in coding agents
SecProbe estimated agent ability with up to 29.5% fewer tasks while finding that no evaluated model solved more than 28.33%.
Top University
Xiaonan Luo · Yue Huang · Kehan Guo · Ping He · Chuan Zou · Chujie Gao · +6 more
University of Notre Dame · Bake AI · Vanderbilt University · University of Pennsylvania · LMU Munich
Research Digest··2 min read
Thread:Agent Security & Attacks
The authors developed SecProbe, an evaluation framework that adaptively selects or synthesizes repository-scale vulnerability-repair tasks using Item Response Theory, a statistical method for matching test questions to ability.
Why this paper
From Inria and 7 others · Part of Agent Security & Attacks, now 45 papers
In one line
SecProbe adaptively evaluates coding agents on cybersecurity vulnerabilities, requiring up to 29.5% fewer tasks while maintaining comparable ability estimates.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§