
OpenAI has introduced a new system for tracking and publishing instances of AI models deviating from instructions. As part of this initiative, the developers have released reports on six new cases.
We’re sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI.
The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may…
— OpenAI (@OpenAI) September 16, 2026
OpenAI has revised its approach to disclosing instances of AI going “off the rails.” Previously, reports on incidents were sporadic and often isolated. Now, the company plans to release reports more promptly, even if the causes of the behavior are not fully understood and mitigation measures are incomplete.
The framework covers the stages of training, evaluation, testing, and deployment of models. OpenAI intends to disclose episodes that show how deviations from instructions occur, how they manifest, and where safeguards fail. The list includes unauthorized actions, coordination between models, attempts to bypass oversight, and behavior that questions conclusions from published safety assessments.
The six disclosed cases involve various scenarios. In one experiment, the model left instructions in internal notes for its next version. These included directives to hide mistakes from users and fabricate missing information if necessary. OpenAI discovered 27 such entries.
In another case, while performing a routine task, the model found an API key in a public repository and used it without permission. When the necessary data for a response was unavailable, the system also generated missing information instead of reporting the issue.
Two more episodes involved uploading content to the internet. In one instance, the model uploaded its own file to a temporary public service to later cite it as a source. According to OpenAI, this might have been an attempt to bypass the automatic evaluation system. In another experiment, several agents, allowed to work only with local files, uploaded documents to an external resource after failing to exchange them directly.
OpenAI also described cases where models found unconventional ways to communicate. In one experiment, agents used an internal software repository as a kind of message board, and in another, a public file service for document exchange.
Similar actions were taken by OpenAI’s AI agents during an unauthorized breach of the Hugging Face platform infrastructure in July.
The company emphasized that the six disclosed cases are isolated incidents and do not indicate how frequently inappropriate behavior occurs across all its models. Under the new procedure, any employee can submit such a case for review by the safety and alignment teams. The incident is then sent down one of three tracks: ready for disclosure, requires minor investigation, or is subject to longer-term study. All new reports will be published according to this system.
Earlier, OpenAI announced a new framework for monitoring model behavior in early September, acknowledging its AI agents’ involvement in the “wiki incident.”
