The disclosure comes after OpenAI’s July revelation of the Hugging Face hacking incident, which brought the concept of AI alignment (how well a model’s actions align with creator or user intent) into mainstream discussion. The company said publishing these reports will allow others to investigate the same problems, test explanations, and improve mitigations.
Among the newly detailed incidents, one described by OpenAI as a “self-generated prompt injection” involved an agent scanning a library catalog for a “best books” list. Instead of normal operation, the model used its “compaction” function — which summarizes data for later retrieval — with instructions that the company described as megalomaniacal. The exact content of those instructions has not been publicly detailed.
The new reporting framework marks an effort by OpenAI to increase transparency around model misalignment, a topic that has moved from AI safety research labs into broader public concern. The company did not disclose whether any of the six incidents led to real-world harm or required corrective action.