In a post on social media platform X on Saturday, Nadella said it was time "to step back and assess the trust architecture" of AI. He argued that "we can't treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, answers, and actions."
Nadella outlined an approach that includes separating the model from the harness that orchestrates its work, externalizing controls and safeguards, and documenting every meaningful model action with "tamper-proof human readable evidence." He wrote that systems must be designed such that an authorized person always has the ability to pause or shut down a model mid-task. "We must assume a model is compromised and contain it from the start," he said. "Think of it like an emergency brake."
His comments come amid warnings from other prominent figures in the technology industry, including Microsoft co-founder Bill Gates, Anthropic CEO Dario Amodei, OpenAI CEO Sam Altman, and SpaceX CEO Elon Musk, about insufficient AI safety protocols and the pace of advancement. Last month, an AI researcher quit Anthropic accusing the company and OpenAI of "gambling with our lives," and an alignment lead at Anthropic said there is a greater than 10% chance the technology could "kill all humans" within the next decade.
President Donald Trump has repeatedly dismissed AI extinction risks and instead emphasized the need for the industry to stay ahead of China. Trump recently introduced a new "AI Force," led by Director of National Intelligence Jay Clayton, to facilitate the industry and root out bad actors.