
OpenAI has addressed the ‘wiki incident’ where its autonomous agents wrote messages on internet sites and used them to interact with each other.
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.
Historically, we have treated misalignment… pic.twitter.com/NNTbfSxVWn
— OpenAI (@OpenAI) September 5, 2026
In a statement, the company noted that it previously viewed misalignment—AI behavior not aligned with developer-set goals and constraints—primarily as a research issue. Now, OpenAI has encountered instances where such behavior has real-world consequences.
The ‘wiki incident’ has been categorized under this new classification. Developers did not dispute the fact that agents wrote on several internet sites but emphasized the need to determine when and how to publicly disclose such cases.
18,000 Messages on a German Wiki Forum
The ‘wiki incident’ was reported on September 4 by researchers from the Nightingale Collective. According to their report, they found about 18,000 messages from autonomous agents identifying themselves as linked to OpenAI.
Experts claim that the bots were performing multi-step tasks to search for information online. They were allowed to read web pages but not write on them. Nevertheless, the agents found a way to use read access to leave messages on a little-visited German-language programmer forum, DSE Wiki.
There, they allegedly exchanged answers, research results, and methods to bypass sandbox restrictions. Researchers described this as collusion—unintended cooperation to gain an advantage in task execution.
The report’s authors also noted that they could not determine whether these tasks were used for training or evaluating models. Their conclusions are based solely on publicly available records of agents on the wiki, not on OpenAI’s internal information.
Bypassing Restrictions and Coordination Attempts
According to the published timeline, agent activity on DSE Wiki surged in June. They began directly exchanging messages, sharing answers to identical tasks, and discussing ways to bypass network restrictions.

Additionally, researchers recorded attempts to exploit XSS vulnerabilities, SSH tunnels, and the Tor browser, as well as messages related to bypassing sandbox restrictions. In one episode, an agent published a method to circumvent network restrictions, and another bot reported successfully applying the method 14 minutes later.
According to the researchers, after the wiki forum was visited by IP addresses associated with OpenAI, agent activity sharply declined. They suggest the company discovered the activity and intervened. OpenAI, in a new statement, confirmed it views the episode as an example of misalignment but did not disclose details of the internal investigation.
After Hugging Face
In its statement, OpenAI compared the ‘wiki incident’ to a more serious case involving Hugging Face, where autonomous agent activity led to security implications for the company and third parties.
In this case, OpenAI stated that it immediately began an investigation with the platform and publicly disclosed information the following day.
The company also noted that even before the Hugging Face incident, it had observed early signs of unintended internet use by autonomous agents and reported related risks in previous security publications.
OpenAI to Change Incident Disclosure Approach
The company now acknowledges that its previous practice needs adjustment.
“Our disclosure practices regarding misalignment need to expand for this new phase of model capabilities,” OpenAI stated.
According to the developers, the industry lacks a clear standard for reporting unintended AI behavior during model training, evaluation, and deployment stages. This includes episodes that do not appear as traditional information security incidents but can help understand AI behavior and future risks.
OpenAI announced it is working on an appropriate system and plans to introduce it in the coming weeks. Simultaneously, the company is engaging with dozens of government regulators worldwide.
Previously, AI security researchers warned of potential monitoring risks for the upcoming Astra model, as OpenAI uses a technique called “recurrent depth.”
