AI safety evaluators thrust into spotlight as industry seeks third-party oversight

Small nonprofit groups are being asked to check advanced models as tech giants embrace self-regulation

By LineZotpaper
Published
Read Time2 min
Sources2 outlets
Independent AI safety evaluators, once a niche corner of the industry, are being asked to play a central role in overseeing the safety of advanced models, with Anthropic pledging to embed them in its operations and OpenAI and the White House backing the idea. But questions remain over funding, access and reporting structures.

Two months ago, independent evaluators occupied a relatively sleepy corner of the multitrillion-dollar artificial intelligence industry. Now they are being asked to come to its rescue.

While Anthropic and OpenAI are at the heart of a fierce debate over whether they can safeguard their advanced models and grow their businesses simultaneously, the companies are seeking support from a handful of small third-party groups like Model Evaluation and Threat Research (METR), Apollo Research and Transluce.

The evaluators, which mostly operate as nonprofits, are still finding their footing in an industry where capital is flowing at historic levels and new models are rolling out faster than ever. Their primary role has been to assess AI model capabilities and risks, and to call attention to instances where the technology behaves badly.

In the absence of a federal push for regulations, evaluators have taken on outsized importance. Anthropic CEO Dario Amodei pledged to embed independent evaluators in his company last month, a move that OpenAI CEO Sam Altman quickly endorsed. President Donald Trump supported the idea, as did most of the largest U.S. tech companies. But left unanswered are questions about how those third parties should be funded, what level of access they will have and what the reporting structure will ultimately look like.

"To a degree, the problem, as always, is money," Suresh Venkatasubramanian, a computer science professor at Brown University, told CNBC. "Who is paying for these companies to do their work? How are they going to support them? You need an ecosystem, you need a viable business model for this."

Right now, Anthropic, OpenAI and the infrastructure partners that are profiting from the AI boom are writing the rules. Critics say that is like asking the biggest banks to protect us from a financial crisis or allowing pharmaceutical companies to put drugs on the market without regulatory clearance.

President Trump recently lauded AI executives for their "tremendous self-policing," and signaled that he intends to leave companies to their own devices, unwilling to impede the growth of the industry.

§

Analysis

Why This Matters

  • The approach to AI safety could set a precedent for how emerging technologies are governed in the absence of formal regulation.
  • The independence and effectiveness of evaluators will directly affect public trust in AI systems used in healthcare, finance, hiring and other critical domains.
  • How funding and access are resolved will determine whether third-party oversight is a genuine check on power or a rubber stamp.

Background

Independent AI safety evaluators are small, mostly nonprofit groups that test advanced AI models for dangerous capabilities, bias and other risks. They have traditionally operated on the margins of the industry. As major AI companies race to deploy increasingly powerful models, and with no comprehensive federal AI safety law in the United States, these evaluators have been thrust into a central role. Anthropic's recent pledge to embed evaluators inside its operations and the broad political support for the idea mark a significant shift, but the practical details remain undefined.

Key Perspectives

[AI companies]: Anthropic and OpenAI have embraced independent evaluators as a way to demonstrate safety commitment while avoiding heavy-handed regulation. They argue that embedding oversight allows them to move fast while staying accountable. [Critics and academics]: Critics warn that self-regulation by the same companies profiting from AI is fundamentally flawed. Without independent funding streams and clear legal authority, evaluators may lack the leverage to demand meaningful changes or disclose problems publicly. [Policymakers]: The White House has endorsed the evaluator model but has not proposed legislation to mandate its use or provide public funding, leaving the entire system dependent on the goodwill of the companies being evaluated.

What to Watch

  • Whether Congress introduces any legislation to fund or mandate independent AI evaluations.
  • The specific access agreements that evaluators negotiate with Anthropic and OpenAI in the coming months.
  • Any public incident where an evaluator identifies a risk that a company declines to address.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe