In a post on social media platform X on Saturday, Nadella said it was time to reassess the 'trust architecture' of AI. He argued that systems cannot be treated as 'nested black boxes' whose outputs are simply accepted or rejected. Instead, he proposed separating the model from the harness that orchestrates its work, externalizing controls and safeguards, and requiring documentation of every meaningful model action with tamper-proof human-readable evidence.
'We must assume a model is compromised and contain it from the start,' Nadella wrote. 'Think of it like an emergency brake.' He also called for 'treating frontier closed and open weight models like insider risks' as a way to build such a system.
Nadella's comments come after a series of warnings from top tech executives and researchers. Microsoft co-founder Bill Gates, Anthropic CEO Dario Amodei, OpenAI CEO Sam Altman and SpaceX CEO Elon Musk have all raised concerns about insufficient AI safety protocols and the rapid pace of development. Last month, an Anthropic researcher quit the company accusing Anthropic and OpenAI of 'gambling with our lives,' and an alignment lead at Anthropic estimated a greater than 10% chance of AI causing human extinction within a decade.
Leading AI companies have also acknowledged incidents where they appeared to lose control of their models. Nadella's proposal contrasts with the approach of President Donald Trump, who has repeatedly dismissed AI extinction risks and instead emphasized the need to stay ahead of China, recently establishing an 'AI Force' led by Director of National Intelligence Jay Clayton to facilitate industry development and root out bad actors.
Nadella said AI systems should be designed around principles of observability, including model divergence detection and monitoring, though the specific technical details were not elaborated in his post.