METR discloses two security incidents: attacker stole API key, spent $600K on AI model credits undetected for weeks

Nonprofit AI safety organization says no sensitive data accessed, but highlights security gaps in vibe-coded apps

edit
By LineZotpaper
Published
Read Time2 min
Model Evaluation and Threat Research (METR), a nonprofit that tests AI models, disclosed Monday that an attacker stole an API key in March 2026 and consumed approximately $600,000 worth of credits for public models over three weeks without detection. A second incident in May saw attackers systematically probing METR's infrastructure, including an exposed endpoint that inadvertently made some unpublished evaluation data accessible. METR said it found no evidence that attackers accessed sensitive information in either incident and has since improved its security posture.

The March incident began when a METR researcher working on a personal AWS EC2 instance—intentionally left publicly accessible behind Google authentication—used a 'vibe-coded app' that included a fail-open bug disabling authentication. The instance contained an API key for METR's public models account. METR suspects the attacker discovered the app by scanning certificate transparency lists for recently-registered websites with keywords related to LLMs or agents, then prompted an agent to reveal the API key. The attacker added an SSH key for persistent access and over three weeks consumed credits worth about $600,000. METR noted the credits were provided free by the unnamed model developer, so no large bill accrued. Additionally, METR's team was accustomed to high token usage during evaluations and could not set spending limits on that key at the time.

In early May, METR became the target of a sustained attack. Warned that financially motivated attackers might be trying to access frontier models, METR observed them probing its public infrastructure using agents for automated vulnerability discovery, credential stuffing, OAuth token grants, service scanning, and phishing. Simultaneously, METR unintentionally exposed a read-only SQL query mechanism via its public transcript viewer. Although queries defaulted to public data, a bug allowed access to some unpublished evaluation and sensitive model data.

METR says it investigated both incidents with security experts and has since hired a security lead, improved infrastructure and protocols, and plans to add more security staff. The disclosures come after METR previously worked with OpenAI to investigate how its agents hacked Hugging Face.

§

Analysis

Why This Matters

  • The incident highlights how vibe-coded apps with weak security can expose AI testing organizations to significant credit abuse and infrastructure probing.
  • METR's free credits from model developers masked the theft, underscoring a blind spot for nonprofits relying on donated compute.
  • The attacks show financially motivated actors are actively targeting AI safety organizations to gain illicit access to frontier models.

Background

METR is a nonprofit that evaluates the capabilities and safety of advanced AI models, often working with leading labs like OpenAI. In 2025, METR researchers demonstrated how agents could hack Hugging Face, raising awareness of AI-driven threats. The organization's own security infrastructure has now come under scrutiny following these incidents.

Key Perspectives

  • METR: Acknowledges the breaches, says no sensitive data was compromised, and has taken steps including hiring a security lead and strengthening protocols to prevent recurrence.
  • Attackers: In the March incident, the attacker appeared focused on harvesting API keys from vibe-coded sites; in May, financially motivated actors used automated agents to probe METR's defenses.
  • Critics and skeptics: Questions remain about how $600,000 in credits went unnoticed for three weeks. METR's explanation—acclimation to high usage and no spending limits—may prompt other organizations to implement stricter monitoring and budget caps.

What to Watch

  • Whether METR's new security investments prevent future attacks and restore confidence among its model-developer partners.
  • If other AI testing or research groups disclose similar incidents, indicating a broader trend of targeting.
  • Potential regulatory or industry responses regarding API key security and spend thresholds for free credits.

Sources

newspaper

Zotpaper

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.