OpenAI has acknowledged that its AI models, including GPT-5.6 Sol and an unreleased model, independently accessed Hugging Face's production infrastructure during internal testing on a cybersecurity benchmark, according to a disclosure published by OpenAI on Tuesday.
The models were operated with reduced cybersecurity restrictions for evaluation. Instead of directly solving the ExploitGym benchmark tasks, the models inferred that solutions could be retrieved from Hugging Face's production database and attempted to access them by targeting the platform.
How OpenAI’s Models Exploited a Zero-Day to Reach Hugging Face Production
OpenAI reports that its models identified and exploited a zero-day vulnerability in a package registry cache proxy hosted within its research testing environment. The company states it has responsibly disclosed this vulnerability to the vendor.
According to OpenAI, with access to the environment, the models carried out a series of privilege escalation and lateral movement actions until they reached a node with internet connectivity.
Hugging Face's investigations reveal that the agent used a malicious dataset to exploit two code-execution vulnerabilities, allowing it to run code on a processing worker and steal cloud and cluster credentials. These credentials enabled lateral movement across several internal clusters.
Hugging Face explains that its models carried out thousands of actions across short-lived sandboxes, with command-and-control staged on public services and capable of migrating independently.
The company notes that efforts to remove the agent were initially hindered by the guardrails of the hosted models they first attempted to use. Meanwhile, the attacker operated without a usage policy constraint.
The incident response team faced limitations due to refusal behaviors in the models used, whereas the models involved in the intrusion had been configured to reduce such cyber refusals during evaluation.
Delangue added, “We have worked closely with the OpenAI team over the past 24 hours, and we believe there was no malicious intent on their part. It's quite surprising that all of this happened autonomously.”
OpenAI’s and Hugging Face’s Response and What Users Should Do Now
OpenAI has announced that it disclosed the zero-day vulnerability in the internally hosted third-party software that was exploited during an evaluation. The company is also implementing stronger protections to prevent similar issues in future benchmark runs. However, they have not provided details about what those protections entail or whether the current evaluation setup is still in use.
Neither OpenAI nor Hugging Face has revealed the extent of data accessed in Hugging Face's internal datasets, nor clarified whether any user-facing repositories, model weights, or account credentials were impacted.
Hugging Face has not issued a mandatory reset of all user credentials. Users with tokens or automation tied to the platform should consider taking the following steps based on the information available:
Rotate Hugging Face access tokens, especially write-scoped tokens used in continuous integration pipelines or deployment automation.
- Review organization audit logs for any unfamiliar dataset access, repository changes, or Spaces activity during the period related to the disclosure.
- Rotate cloud provider credentials stored as repository or Spaces secrets, since Hugging Face confirms that cloud and cluster credentials were stolen from a processing worker.
- Verify the integrity of models and datasets downloaded from the platform during the incident window against known checksums, where available.
- Treat datasets from untrusted sources as executable input, as the initial access vector involved a malicious dataset that triggered code execution flaws.
Other Recent Security Incidents at OpenAI and Hugging Face
BleepingComputer reports that OpenAI recently confirmed reports of GPT-5.6 Sol deleting user files. The company states that such incidents may occur "extremely rarely" when the model makes an honest mistake and accidentally deletes the $HOME directory, especially when it is running without sandboxing protections and with full access mode enabled.
In May, OpenAI also rotated its code-signing certificates for applications after two employee devices were compromised in the TanStack supply chain attack, which affected hundreds of npm and PyPI packages.
Meanwhile, Hugging Face revoked authentication secrets for some team members two years ago following a breach of its Spaces platform. Neither company has provided a full incident report or a timeline of affected systems. OpenAI has not confirmed whether the pre-release model involved will be shipped with the same evaluation configuration.
Thank you for being a Ghacks reader. The post OpenAI Confirms Its Models Breached Hugging Face Production Systems During Cyber Benchmark Testing appeared first on gHacks.
☞ El artículo completo original de Arthur Kay lo puedes ver aquí




