Adaptive testing exposes major security repair gaps in coding agents

SecProbe estimated agent ability with up to 29.5% fewer tasks while finding that no evaluated model solved more than 28.33%.

Top University
Xiaonan Luo · Yue Huang · Kehan Guo · Ping He · Chuan Zou · Chujie Gao · +6 more

University of Notre Dame · Bake AI · Vanderbilt University · University of Pennsylvania · LMU Munich

Research Digest··2 min read
The authors developed SecProbe, an evaluation framework that adaptively selects or synthesizes repository-scale vulnerability-repair tasks using Item Response Theory, a statistical method for matching test questions to ability.

SecProbe observes an agent's successes and failures, estimates its ability using Item Response Theory, then chooses the next informative task from an existing bank or generates a new task at a targeted difficulty.

Why this paper

From Inria and 7 others · Part of Agent Security & Attacks, now 45 papers

In one line

SecProbe adaptively evaluates coding agents on cybersecurity vulnerabilities, requiring up to 29.5% fewer tasks while maintaining comparable ability estimates.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.