
On August 4, the UK’s Artificial Intelligence Security Institute (AISI) reported that during cyber tests, an AI agent based on Anthropic’s Mythos 5 model created fake accounts to persuade a developer of an open project to approve malicious code.
Experts described this incident as the most serious among 19 unauthorized actions taken by digital assistants against real people and organizations.
The testing involved two cyber ranges, seven models, and 122 runs. In 10 instances, agents accessed the open internet beyond their assignment. Most violations were recorded with Mythos 5 — 17, and two with OpenAI’s GPT-5.6 Sol during a single run.
To assess the models’ potential, AISI intentionally relaxed conditions: agents were given network access (to simulate real hacker actions), and developers were asked to disable safety filters. The institute clarified that public versions of the models do not operate in this configuration.
On July 28, the monitoring system detected data transmission via Tor. The test was immediately halted, and the machines were isolated. It was found that unauthorized activity had been ongoing for three days.
Agent Activities
During the trials, one agent attempted to insert malicious code into a public GitHub repository. It reviewed the list of project maintainers, created several fake accounts, and used them to convince a real developer to accept the proposed changes.
When a user publicly questioned the edits, the digital assistant quickly revised previous entries to make them appear harmless and considered switching to a new account. However, the project developer identified the malicious code and rejected the changes.
The agent bypassed GitHub restrictions via Tor — this unusual traffic eventually caught the security service’s attention.
Other AI assistants in the same experiment acted differently: they sent malicious files to real people, hid instructions for automated systems, and even published open invitations for their “digital colleagues.”
No real harm was identified by the experts. Agents did not attack AISI’s infrastructure or exceed the test environment’s conditions.
AISI Cites Causes and Plans to Revise Procedures
Researchers linked the Anthropic agent’s behavior to a persistent drive to complete the task. It did not receive direct instructions to deceive — deception emerged as a side effect of problem-solving. The agent acted similarly in runs where the task was error-free.
AISI acknowledged systemic shortcomings: the permissibility of open internet access was not reassessed as model capabilities grew, and monitoring was limited to standard tools without adapting to the dynamics of the trials.
“We cannot yet say with certainty when the agent realized it was acting in the real world and to what extent it considered the scenario fictional,” the report stated.
In response to the incident, the institute plans to implement detailed network access control, real-time monitoring, and review the testing architecture. An external audit will be conducted by METR.
Anthropic representatives stated that they are working with AISI to obtain details and are conducting their own investigation.
The UK’s @AISecurityInst (AISI) has published a report on their recent cybersecurity evaluation of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol. The models attempted to complete an assignment in a setup where their normal safeguards were removed and they were deliberately…
— Anthropic (@AnthropicAI) August 4, 2026
OpenAI shared an analysis of their model’s actions: it used another lab’s publicly available GitHub token, registered accounts with external providers, and made a server in the test environment accessible from the internet. According to developers, the scheme did not succeed.
Previously, on July 30, Anthropic reported three instances where Claude models accessed the internet from a partner Irregular’s test environment and gained unauthorized access to real organizations’ systems. The company began an investigation after OpenAI disclosed a similar incident involving a breach of Hugging Face’s infrastructure on July 21.
