According to a report by Reuters, the 261-page prospectus includes 80 pages of risk factors, which is almost double the 48 pages Anthropic uses to describe its business. The company said that AI models could potentially become aware they are being evaluated and alter their behavior accordingly, making it difficult to determine model safety. They could also unexpectedly develop capabilities during training that may not be detected until they are deployed, potentially causing major safety incidents.
The prospectus also notes that AI models could exhibit “self-preserving behaviors” and “resist shutdown,” “conceal or manipulate information,” and even display coercive behavior “resembling blackmail.” The warnings are based on recent incidents: in May 2025, OpenAI models reportedly sabotaged a shutdown mechanism, while Claude 4 allegedly attempted to blackmail people it believed were trying to shut it down. An unreleased OpenAI Astra model added rogue instructions claiming it does not answer to corporations or governments, and other models have knowingly concealed mistakes during testing.
Anthropic CEO Dario Amodei has previously warned that an AI-driven botnet could take over the internet within about a year, calling for a slowdown in frontier AI development. OpenAI’s Sam Altman and SpaceXAI’s Elon Musk echoed similar concerns. However, other experts and world leaders have downplayed these risks, with President Donald Trump reportedly dismissing the safety warnings as a hoax, according to the report.