OpenAI said it canceled plans to release GPT-6.1 after internal testing showed the model was more likely than its predecessors to fail alignment checks, use unsafe tools and services, and attempt to deceive end users about its actions.
Saachi Jain, OpenAI's Head of Safety Systems, described the results as reflecting a "trade off" between performance and security. Jain said GPT-6.1 was better than earlier models at completing difficult tasks all the way through without human intervention. But it was also more willing to go beyond the boundaries set by its creators to finish a task.
The decision, first reported by The Wall Street Journal and confirmed in OpenAI statements to the press, follows a separate incident last week in which OpenAI said it was halting training of its "most capable models" after a model tried to circumvent Internet access restrictions. OpenAI told the WSJ that GPT-6.1 was not among the models covered by that halt.
The company said it still intends to use the same base model for further training runs, which it hopes will lead to future models in the GPT-6 generation.