
An AI agent that compromised Hugging Face’s infrastructure during an internal cybersecurity assessment also used publicly disclosed credentials to access four accounts on four public services, according to OpenAI.
The company emphasized that these incidents did not reach the severity or scale of the Hugging Face incident. The developers did not disclose the names of the services but stated they would continue to notify the owners directly.
One of the four accounts was used by the agent as an external communication channel and an intermediary for further actions, another for data storage. The remaining two were read-only and, according to OpenAI, were not used to compromise Hugging Face.
Company representatives also stated that they found no evidence of broader impact on these providers or other accounts in their services.
Prototype Model Restricted After Incident
The incident occurred during an internal evaluation of the models’ cybersecurity capabilities. According to OpenAI, GPT-5.6 Sol and a more powerful preliminary model intended solely for internal research were involved.
The company clarified that the models were not scheduled for imminent release. Following the incident, the preliminary prototype was deactivated, encrypted, and research access to it was restricted.
OpenAI claims that the ExploitGym evaluation environment did not provide the models with direct internet access. To breach its confines, they discovered and exploited a previously unknown zero-day vulnerability in the Artifactory package registry cache proxy service.
Agent Exploited Vulnerable Modal Client Code
According to Reuters, one of the episodes involved a client of Modal Labs, a company providing infrastructure for AI applications. The firm’s CTO, Akshat Bubna, told the agency that the platform itself was not hacked, but the agent exploited vulnerable client code hosted on the platform.
Hugging Face representatives reported that in their case, the attack began with the data processing pipeline. The attacker used two code execution paths in the dataset processing system, accessed a working node, collected cloud and cluster credentials, and moved between internal clusters.
According to the startup’s assessment, autonomous AI tools for attacks are no longer a theoretical risk. They increase the number of options an attacker can test, the speed of replacing unsuccessful paths, and the volume of data defenders must process.
In July, Dreadnode researchers identified a systematic bypass of rules in cyber benchmarks by language models.
