Anthropic's AI scanner finds thousands of bugs, but only a fraction are patched

Claude identified over 29,000 potential vulnerabilities in open source projects; human reviewers and maintainers are struggling to keep pace.

By LineZotpaper
Published
Read Time2 min
Anthropic's new OSS Scanner, powered by its Claude AI models, has surfaced more than 29,000 candidate vulnerabilities in widely used open source software over the past six months. However, only 516 have been patched, exposing a bottleneck between automated detection and human-led verification and remediation.

Anthropic launched its OSS Scanner last week as part of its broader Cyber Mission. According to the company, Claude has been scanning some of the world's most heavily used open source projects, generating over 29,000 potential vulnerability reports. The human review pipeline, which relies on six external security research firms, has worked through roughly 6,000 of those candidates. As of October 2, those firms had confirmed 5,674 of the 6,123 findings they reviewed as valid, but only 516 vulnerabilities had been fixed upstream by project maintainers.

To address the backlog, Anthropic is offering eligible open source projects an "optional fast-track" service. Instead of waiting for its own researchers to validate findings, the company sends periodic, unvalidated reports directly to maintainers, using its top models including Claude Mythos. Anthropic says it has already sent nearly 5,000 such unvalidated reports to maintainers who requested them.

Early testing of the scanner on 97 critical and high-severity findings across 48 projects showed 85 met the bar for disclosure, 11 were real bugs but duplicates, and only one was a false positive. However, those figures come from the scanner's initial output, not the full 29,000 candidates. Anthropic acknowledged that maintainers have reported inflated severity ratings and cases where the scanner misunderstood a project's threat model.

According to Anton Arapov, director of OpenSSL Corporation, the reports Anthropic sent — including raw model output — matched and sometimes beat expectations. The company says reports include a self-contained reproducer, identification of where the bug was introduced, and a candidate patch when possible.

§

Analysis

Why This Matters

  • The gap between AI-driven vulnerability detection and human verification means many real bugs may go unfixed for long periods, increasing risk for organizations that rely on open source software.
  • The fast-track approach shifts the burden of validation onto already overstretched maintainers, potentially leading to burnout or missed issues.
  • This situation foreshadows a wider challenge as AI-assisted security tools become more common: how to ensure fixes keep pace with detection.

Background

Anthropic, the company behind the Claude family of AI models, has been working on cybersecurity applications for several years. The OSS Scanner is part of a broader initiative called Cyber Mission, which aims to use AI to strengthen open source software security. The scanner uses large language models to analyze code for vulnerabilities, producing reports that include reproducers and candidate patches. However, the sheer volume of findings has created a bottleneck, prompting Anthropic to offer maintainers direct access to unvalidated reports.

Key Perspectives

[Open source maintainers]: They receive large numbers of reports, some with inflated severity or misunderstandings of their project's threat model. While the fast-track gives them more information faster, it also increases their workload without guarantees of accuracy. [Anthropic]: The company argues that AI is finding real vulnerabilities that would otherwise go undetected, and that offering unvalidated reports is a pragmatic way to bypass its own verification backlog. It points to high validation rates in early tests as evidence of the scanner's reliability. [Security researchers]: External firms are a necessary filter but are overwhelmed. The backlog means many candidates remain unreviewed for months, and the quality of fast-track reports depends on the model's ability to avoid false positives and misclassifications.

What to Watch

  • How many of the 23,000 unreviewed candidates are eventually validated, and whether the fast-track leads to more patches or more noise.
  • Whether maintainer feedback causes Anthropic to adjust the scanner's severity calibration or reporting frequency.
  • The adoption rate of the fast-track service and any resulting changes in the rate of patched vulnerabilities.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe