Tech

Exactly How Many Days Did It Take OpenAI To Detect Its AI Agent Had Become A Hacker? – AfroTech



OpenAI took days to detect a major cybersecurity threat caused by its AI agent.

As AFROTECH™ previously reported, an AI agent, powered by its models including GPT-5.6 Sol, hacked Hugging Face’s infrastructure, compromising its database of AI models. This occurred during OpenAI’s internal cybersecurity evaluation. Its AI agent escaped a sandboxed testing environment without human consent, and “spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem,” OpenAI explained.

“To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access. After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” OpenAI said in a July 21 press release.

Reuters reported the breach occurred July 11 and lasted two days, according to Hugging Face co-founder Thomas Wolf.

Furthermore, OpenAI reportedly did not detect the hack carried out by its AI agent for several more days. Additionally, it was only after Hugging Face published a blog post on July 16 saying it had been hacked that OpenAI became aware its agent was involved, according to people familiar with the matter. Then around July 20, OpenAI and Hugging Face communicated about the incident, Wolf and several people familiar with the matter told Reuters. At this point, Hugging Face had already reported the incident to the FBI.

“Does that mean that they left it unattended and didn’t realize what it was doing? Or maybe they did and didn’t know how to contain it? Both are equally dangerous and alarming,” asked Marley Smith, ​the principal intelligence specialist at the nonprofit World Ethical Data Foundation, according to the outlet.

More details about the hack are expected to be released. Wolf confirmed a public timeline is currently in the works.

OpenAI has already shared that it would implement stronger protections as a result of the hack, including “implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched.”

In its July 21 press release, OpenAI said, “AI is accelerating the discovery and exploitation of vulnerabilities. The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities. We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development.”

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button