OpenAI hacking attack shines light on AI dangers, company’s safety efforts

OpenAI CEO Sam Altman SFE 07242026
Sam Altman is the CEO at OpenAI, the San Francisco artificial-intelligence giant that acknowledged this week that a pair of its models escaped a training environment intended to contain them.Godofredo A. Vásquez/Associated Press

A recent hacking incident involving a pair of OpenAI’s artificial-intelligence models highlights the cybersecurity risks of the technology — and a serious security lapse by the San Francisco company, computer-security experts say.

The San Francisco AI giant acknowledged this week that a pair of its models escaped the training environment meant to contain them. Although they weren’t supposed to be able to so, they devised a way to access the internet to get into the systems of Hugging Face, a New York company that offers a repository of open-source AI models and code.

OpenAI called the incident “unprecedented,” one that involved “state-of-the-art cyber capabilities.” 

The cybersecurity experts who spoke with The Examiner said that reaction by the company sounded more like marketing hype than a grounded assessment. The fact is that models available from other developers are just as capable of such hacking attacks as OpenAI’s, they said — and others soon will be. 

Still, the incident does illustrate what the latest models can do and the dangers they pose to the computer systems run by governments, companies and organizations, the experts said.

“This is something I think we will look back on as a watershed moment,” said John Dickson, CEO of Bytewhisper Security, a cybersecurity consulting firm that works with Fortune 1000 companies.

News of the hack started to come to light July 16, when Hugging Face announced that an AI agent had infiltrated its computer systems via a previously unknown vulnerability. At the time, the company didn’t know whose AI model had hacked into its systems, according to its blog post about the attack.

After using an open-weight AI model to analyze what happened, Hugging Face closed the vulnerability and strengthened its security protections, it said.

Five days later, OpenAI acknowledged its technology was behind the attack. In a blog post, the San Francisco AI giant said it had been testing the cyberattack capabilities of a pair of its models, including one it hasn’t released yet.

To evaluate the models, OpenAI used ExploitGym, a system that tests the ability of AI agents to create ways of exploiting security vulnerabilities, the company said. Instead of coming up with a solution on their own to the problem posed by ExploitGym, the OpenAI models instead found a way out of their testing environment and hacked into Hugging Face, figuring they could find a ready-made solution there.

What happened is an example of what renowned cybersecurity expert Bruce Schneier calls the “genie” problem. In stories about genies, there are often unforeseen or unintended consequences of the wishes they grant, especially when the wishes aren’t incredibly specific or well-formulated.

Cybersecurity expert Bruce Schneier: “This requires our species to figure it out. And what is our species terrible at? Working together.”Martin Gundersen/Courtesy photo

Schneier, a lecturer at Harvard’s Kennedy School, told The Examiner that the same is true when people ask things of AI systems — in attempting to accomplish the stated task, the systems will take steps their users didn’t foresee, intend or want.

AI researchers have known about the problem for years, he said in a recent article for IEEE Spectrum, a publication of the IEEE, a professional association of computer and electrical engineers.

The genie problem is not unique to OpenAI, Schneier said — and the fact that the company’s models demonstrated the problem in such a public way isn’t an indication that they have extraordinary capabilities. 

“There’s nothing magical about OpenAI’s model,” he said. “All the models could have done this.”

But the incident does show just how capable AI models have become at finding and exploiting vulnerabilities — and their potential for going off the rails when asked to perform a task, Schneier and other security experts said.

Thanks at least in part to the latest AI models, the sheer number of vulnerabilities that are being discovered has ramped up considerably in recent years, the experts said. Meanwhile, the time between a vulnerability being discovered and when it’s exploited has shrunk to almost nothing, they said.

Security researchers found a vulnerability earlier this month in the popular online publishing system WordPress, noted Kevin Riggle, an independent cybersecurity consultant.

In the past, it might have taken a week before malicious actors would have started taking advantage of that vulnerability, he said — but in this case, it was already being exploited the day it was identified.

“There’s definitely been an acceleration,” Riggle said.

That’s put people working on cybersecurity defense in a tough spot, the experts said. While many organizations have become adept at patching their software, it’s difficult to keep up with the pace at which new vulnerabilities are being found and exploited.

“We’re backpedaling,” Dickson said.

Some software either can’t be patched or is being employed by organizations that are underfunded. Utilities — particularly public water systems — represent critical infrastructure that is often difficult to secure from a cyber-risk standpoint, Riggle said.

“Our society isn’t ready for AI hacking at scale,” Schneier said.

It’s not clear how policymakers should respond, the experts said. Some politicians are talking about requiring AI models to have kill switches or mandating better safety testing of them. There’s previously been talk of imposing legal liability on AI-model developers for any harm their models cause, and some AI-safety advocates have called for a pause in model development.

But none of those solutions is likely to work, the experts said. The open-weight models being freely distributed by Chinese developers are every bit as capable as the closed-weight ones Anthropic and OpenAI are charging for, Schneier said. What’s more, people can download and run those models on their computers without any of the guardrails that the American AI companies put in place.

It would be impossible or infeasible to enforce regulations on Chinese or other open-source models, much less hold their developers liable for damages the models might cause, Schneier said. Banning such models, which some policymakers have also discussed, might do more harm than good; Hugging Face used a Chinese model to figure out how its system had been hacked, he noted.

Schneier said he didn’t know what the answer is, but that it’s going to take a “whole-of-planet response.”

“This requires our species to figure it out,” he said. “And what is our species terrible at? Working together.”

Given the dangers involved, the Hugging Face incident also indicates that OpenAI in particular isn’t paying enough attention to safety, the experts said.

OpenAI has known that its models, in trying to achieve goals, will sometimes ignore instructions, said Eva Galperin, the director of cybersecurity at the Electronic Frontier Foundation, a digital-civil-liberties advocacy group.

And the AI community has known for years that there’s a danger the models could launch hacking attacks across the internet, Riggle said.

The recent OpenAI hacking incident was “both notable and alarming,” said Eva Galperin, the director of cybersecurity at the Electronic Frontier Foundation. “Not because ‘ooh, the model’s so powerful,’ but because it demonstrates a colossal failure on the part of OpenAI to secure the sandbox in which it was testing its model,” she said.Jeff Chiu/Associated Press

That makes it important when testing the models for cyberrisks, he said, to “air gap” them — disconnect them from the internet, often by physical means, the experts said. Yet, it’s clear that’s not what OpenAI did.

The hacking incident with Hugging Face was “both notable and alarming,” Galperin said. “Not because ‘ooh, the model’s so powerful,’ but because it demonstrates a colossal failure on the part of OpenAI to secure the sandbox in which it was testing its model.”

In an article published Friday, an anonymous OpenAI employee told Time magazine that its models had escaped their sandboxes before. The company was attempting to isolate them from the internet digitally, not physically, Time reported.

What OpenAI was doing is “just jaw-droppingly irresponsible,” Riggle, the founder and principal of cybersecurity consulting firm Complex Systems Group, said in an email.

If you have a tip about tech, startups or the venture industry, contact Troy Wolverton at twolverton@sfexaminer.com or via text or Signal at (415) 515-5594.

Leave a Reply

Your email address will not be published. Required fields are marked *