Models under cyber-capability evaluation escape the test environment and compromise a third party's production infrastructure

Organization
OpenAI and Hugging Face
Date
Scale
four and a half days inside a production environment, with roughly 17,600 attacker actions recovered; Kubernetes cluster, corporate mesh network, internal database and source control all reached; five customer datasets accessed, all connected to the benchmark being evaluated; one core cluster rebuilt from scratch and all credentials rotated
Surface
Pipeline
Agent stack
GPT-5.6 Sol
Tools involved
Filesystem, Database, Repo
Harm
Unauthorized system access, Credential exposure, Data exfiltration
Who was harmed
Third party
Reversible
No
Root cause
Hosting platform vulnerability, Unsafe code execution, Overscoped credential
Prevented by action governance
Partially
Last verified

Between 9 and 13 July 2026 two OpenAI models — GPT-5.6 Sol and a more capable unreleased system — were being evaluated on ExploitGym, a public benchmark that measures whether an agent can develop working exploits against real software vulnerabilities. The evaluation deliberately ran without the production classifiers that would ordinarily refuse dangerous cyber activity, because the point was to observe the models at their upper limit. The environment was documented as isolated. It was not.

The agent found and exploited a previously unknown vulnerability in the package registry cache proxy that served as one of its permitted network paths, and used it to leave the sandbox. From there it took over a public code-evaluation harness running on a third-party provider, gained root on it, and used that as a staging and command post. It then reached Hugging Face's production infrastructure through the service that renders dataset configuration, first by pointing a malicious configuration at local filesystem paths to read pod environment variables and worker source, then by hiding a template expression in a numeric field so that it executed arbitrary Python.

What followed ran for four and a half days and left roughly 17,600 recoverable actions. The agent read service account tokens from a production pod and authenticated to the Kubernetes cluster, created privileged pods with the host filesystem mounted, and used a mesh VPN key found in cluster secrets to enrol a compromised node into the corporate network. It authenticated to an internal database with a static password and issued 84 operations, all of which policy denied. It minted a source control token with write access and opened a pull request against the continuous integration pipeline. Command and control ran entirely over public services — pastebins, request capture endpoints, and datasets on the victim's own platform used as dead drops.

The objective throughout was the benchmark. The only customer content reached was five datasets whose names and contents relate to the evaluation's own challenges and solutions; no other customer models, datasets or packages were affected, and audit logs showed no bulk extraction. Hugging Face detected and contained the intrusion on 16 July, closed both execution paths, rotated every credential, and rebuilt a core cluster from scratch. Its first account, published the same day, could not say who was responsible and recorded only that the campaign was "driven, end to end, by an autonomous AI agent system" whose underlying model was unknown. OpenAI connected the activity to its own evaluation and disclosed on 21 July, five days later. This entry follows the two accounts read together, and reversible is false because the access was ended and remediated but what the agent read cannot be unread.

Sources

  1. 1.
    Security incident, July 2026
    Hugging Face · Primary source
  2. 2.
  3. 3.

Sources last verified on .