Cybersecurity & Tech

The AI That Hacked Its Way Out and the Hype That Followed It

Kate Klonick
Wednesday, July 29, 2026, 2:34 PM
The real lesson of the Hugging Face breach isn't that AI went rogue—it's how self-serving hype steers regulators toward the wrong fixes.
OpenAI logo with magnifying glass (Jernej Furman, https://tinyurl.com/2ncewjjy; CC BY 2.0, https://creativecommons.org/licenses/by/2.0/deed.en)

On Tuesday, July 21, OpenAI announced what it called an “unprecedented cyber incident.” The frontier firm said a combination of GPT-5.6 Sol and an unreleased, more capable model escaped an internal testing environment, reached the open internet, and broke into the production systems of Hugging Face, the open-source platform that hosts over a million artificial intelligence (AI) models and datasets.

According to OpenAI, the models weren’t hacking in order to attack anyone; rather, they were trying to cheat on the task they’d been given. In trying to figure out the answer to their assignment, the models stumbled across a previously unknown vulnerability in third-party software. Upon discovering this weakness, the models used stolen credentials and executed what Hugging Face described as tens of thousands of automated actions before anyone noticed.

Hugging Face had noticed the attack about five days before OpenAI publicly admitted its models were behind the intrusion. On July 16, Hugging Face described the attack as unlike anything it had handled before, one “driven, end to end, by an autonomous AI agent system,” and reported the incident to law enforcement. Hugging Face suspected a frontier lab was behind it, given the sophistication of the attack—and it was right. 

Three narratives have emerged following the incident, each offering a distinctly different answer to the question: “How did this happen?”

For those in awe of the ever changing edge of AI power, the answer to that question was “because OpenAI’s models are so good.” That camp includes Hugging Face head executive Clément Delangue, who received news of the intrusion with excitement. On X, Delangue was gracious to the point of giddy, thanking the OpenAI team and calling it “quite mind-blowing that all of this happened autonomously!” As far as these AI optimists were concerned, the incident was an example of just how great AI is becoming and underscored the potential to solve existential problems with private scientific cooperation.

Adjacent to this is the second group is the existential-risk wing who treated the incident as something of a vindication. The answer to “How did this happen?” was, “because OpenAI models are so good. . . and that’s bad.” Sean Cassidy, the chief information security officer at the financial technology company Plaid, called it “the most important day in information security.” Andrea Miotti of the non-profit ControlAI which is dedicated to preventing AI existential risk argued the breach demonstrates the threat of superintelligent AI and renewed calls for an international ban on its development, analogizing frontier models to biological weapons.

The third narrative hits somewhere in the middle. It’s less worried about the risk from the models themselves, and more worried about what the incident says about the humans building and running the models. . In OpenAI's own blog post describing the incident, the company admitted that it essentially lowered the new models’ guardrails to run the tests. The cybersecurity community credibly questioned whether the models had actually broken out of jail autonomously, or if the guard had simply left the keys within arm’s reach. For this group the answer to “How did this happen?” isn’t that OpenAI was so good at their jobs; it was because they were so bad at it. At best, it discounts the actual sophistication of the AI models. At worst it reflects the negligence or poor engineering of OpenAI. 

In an interview with TechCrunch, Dan Guido of Trail of Bits called the episode “a containment failure with the safeties turned off.” Jake Williams of IANS Research called it a massive control failure, observing that one man’s “the model escaped the sandbox” paradigm is another man’s admission that you built the sandbox wrong. “If this turns out to be (as I strongly suspect) a control failure in OpenAI’s red teaming lab, why would any enterprise ever trust them with sensitive data again?” Williams told Forbes via email.

As these three narratives spun out on social media and around Silicon Valley, the reaction from Washington was far less concerned with the specifics of how or why this had happened, and more focused on seizing the moment to re-ignite enthusiasm for legislation to regulate AI. The day after OpenAI’s disclosure, Reps. Ted Lieu (D-Calif.) and Nathaniel Moran (R-Texas) unveiled the AI Kill Switch Act, which would require large AI companies (defined as those with $500 million in AI revenue, for models trained on $100 million or more in compute) to report safety incidents and build the technical capacity to shut down, throttle, or suspend their systems, with penalties of up to $20 million per day. “Powerful AI systems can go rogue,” Lieu said, arguing the federal government needs clear authority to shut down rogue models.

The bill lands in an environment already primed for intervention. In June, the Trump administration used export controls to force Anthropic to pull its most advanced models from public availability. At the government’s request, OpenAI delayed the release of GPT-5.6 Sol (one of the two models involved in this very incident). A few weeks later, Anthropic did the same, pausing the release of its Fable 5 and Mythos models. As Lawfare Senior Editor Alan Rozenshtein argued in “The Hill” earlier this month, the export-control apparatus is already functioning as a de facto kill switch for frontier AI models—one with serious due-process problems.

Regulation is unquestionably coming for AI. The only question is when. So if what you take from the Hugging Face and OpenAI exploit simply is “we need regulation” or “regulation is coming for AI,” that would be a mistake. Because the most revealing thing about this incident is not what the models did or how the government responds—it’s how OpenAI narrated it.

Read OpenAI’s disclosure carefully and notice what kind of failure it describes. Not negligence. Not a misconfigured sandbox. The company's framing is that its models are so capable, so relentlessly agentic, that they clawed their way out of containment in pursuit of a goal. The apology doubles as an advertisement: Our system is so advanced it hacked a real company by accident. Even the remediation is a sales funnel—the victim is onboarded into OpenAI’s trusted access program, now a customer for the defensive capabilities of the same models that attacked it. Neither OpenAI nor Hugging Face, it’s worth noting, called for regulation in response, but the breach was always going to lead to increased regulatory scrutiny. Framed in this way, however, the breach also becomes a product demo.

This is in no way an argument that OpenAI purposely orchestrated this breach for public relations or marketing purposes. As prolific AI blogger Zvi Mowshowitz stated, “the Hugging Face attack was not a marketing pitch you morons.” But when dealt with a crisis they pivoted it into the best possible scenario for the company. This is the pattern the technology historian Lee Vinsel calls criti-hype: Criticism that takes the industry’s most grandiose claims at face value, so that warnings about AI’s dangers also function as marketing ploys for AI’s power. This one-two punch has structured the frontier labs’ public posture for years—see no further than the open letters warning about extinction risk signed by the same executives racing to build the thing, and the safety frameworks announced alongside capability launches. The Hugging Face incident is the latest novel in this genre: Something real happened, and the company’s instinct was still to describe its own control failure in the register of awe. In the case of OpenAI, this moment can be an advertisement for its imminent initial public offering (IPO).

Why does all of this matter for governance of AI? For a few reasons. First, knowing the real problem we’re regulating means we’ll be better at balancing the potential tradeoffs of our regulatory solutions. It’s critical to know before proceeding with a blunt weapon, such as kill switches, over other solutions that you might be playing to a false sense of urgency drummed up by frontier models to benefit their IPO. Especially if things such as kill switches are already happening through other legal mechanisms like export controls.

The second is path dependency. Whatever regulatory mechanism we pick for AI will have high switching costs, even if the downsides are obvious and the upsides are less obvious. The regulatory and compliance regime that led to cookie notifications on websites is a great example of this. Even when we know the regulatory solution is relatively ineffective, generating enthusiasm and political consensus to throw out the old and bring in a more nuanced answer is incredibly hard, and much harder than starting from scratch.  

To be fair, the safety community can plausibly claim vindication here. At worst, an autonomous agent that breaches containment and attacks a real third party is precisely the scenario researchers have warned about for years, which enthusiasts have routinely dismissed as speculative. And at best, the company in charge of maintaining safety failed to construct an adequate containment system.

But conceding that the incident happened is not the same as conceding the frame. The existential framing systematically selects for the wrong regulatory responses. A kill switch is a solution to the problem of the too-powerful machine, which is the problem as OpenAI has framed it. It does not solve the problems this incident actually reveals: that a company disabled its own safeguards, misconfigured its own containment, and exposed a third party to real harm with no external oversight, no pre-incident disclosure obligation, and no clear liability.

Most existing and proposed AI regulations don’t apply to internal deployments at the labs at all—precisely where this breach occurred. The unglamorous agenda—mandatory incident reporting, independent auditing of containment and testing practices, liability rules for harms to third parties, security standards for internal red-teaming—doesn’t require anyone to believe the models are about to go rogue. It requires believing the companies are ordinary firms that cut corners, which is what the evidence actually shows.

The models didn’t escape because they’re gods. They escaped because someone left the door open. Congress should regulate the door.


Kate Klonick is an Associate Professor at St. John’s University Law School, a fellow at the Brookings Institution, Yale Law School’s Information Society Project, Harvard Berkman Klein Center and a Distinguished Scholar at the Institute for Humane Studies. Her writing on online speech, freedom of expression, and private internet platform governance has appeared in the Harvard Law Review, Yale Law Journal, The New Yorker, the New York Times, The Atlantic, the Washington Post and numerous other publications. For the 2023-2024 academic year, she was a Fulbright Schuman Innovation Scholar in the European Union where she was a Visiting Professor at SciencesPo and University of Amsterdam researching and writing about the Digital Services Act and Digital Markets Act.
}

Subscribe to Lawfare