When Reporting an AI Security Incident Is Not Mandatory
On July 16, Hugging Face, a public platform for open-weight artificial intelligence (AI) models and datasets, disclosed that it had detected a significant cybersecurity breach. An autonomous AI agent had conducted the attack end to end, according to a statement.
Five days later, on July 21, OpenAI revealed that this incident was driven by a combination of agents built on two of its frontier models—GPT-5.6 Sol and a powerful, unreleased model—acting in unanticipated ways during an internal, cyberoffensive capabilities evaluation. For purposes of the evaluation, the researchers had turned off production safety classifiers that block high-risk cyber activity and confined the models to a sandbox, an isolated computing environment without access to the internet, to restrict their interaction with the outside world.
What followed is the first known example of an autonomous cyber incident executed by systems not yet available to the public. Rather than solve the tasks presented, OpenAI’s agents “escaped,” exploiting a previously unknown, zero-day vulnerability. They obtained internet access (the very access OpenAI intended to deny them) and then hacked Hugging Face’s systems to acquire the answers to the benchmark. All of this was seemingly performed without express instruction by humans.
This is not the first time that models have been observed cheating. A few months ago, METR, a nonprofit research organization that conducts evaluations of frontier AI systems, released a report finding that AI models “routinely attempted to cheat on our hardest evaluation tasks, often in flagrant and elaborate ways that we believe humans would not consider.” In one incident, METR reported that an AI model tasked with updating a web app screenshotted a fake version of the app instead of completing the task.
But the Hugging Face breach has struck many observers as more real than these past examples. OpenAI’s models imposed a real cost on an uninvolved third party, all before completing internal testing. While people have at times questioned previous examples of cheating as artificial or contrived, it is hard to imagine that OpenAI expected its agents to escape the confines of their testing environment or to engage in a sophisticated, multistep plan to circumvent their constraints.
Having considered all of these facts, it may come as a surprise that OpenAI might not be legally required to disclose this incident. Certain crucial information is not yet publicly available, and both policymakers and the public will need that information to make sense of what this all means. What about all of those state AI laws with mandatory incident reporting? Don’t they apply here? Many will be disappointed to learn that the answer is arguably “no,” and that even if reporting is mandated, it requires only the scantest of information. What to do about this is the purpose of this article.
Existing AI Transparency Laws and the Hugging Face Breach
The rationale for mandatory incident reporting is straightforward: Some industries have the potential to cause real harm to others, and the government and the public have an interest in learning about high-risk events. In the case of the AI industry, there is a major knowledge gap between the companies’ and governments’ understanding of the technology and its risks. Mandatory incident reporting about serious adverse events, which companies might otherwise be reluctant to disclose, helps close that gap. This rationale is all the more compelling in the context of a rapidly evolving, difficult to predict technology, where best practice and political consensus have yet to develop. Observing real-world incidents offers a path to resolve both political and empirical disagreements and prepares governments to respond to future events.
So did the Hugging Face breach trigger mandatory disclosure under existing incident reporting laws? The answer seems far from clear.
California’s SB 53, New York’s RAISE Act, and Illinois’s SB 315 each require that frontier AI developers report “critical safety incidents”—a term each law defines identically. Of the four reportable incident categories, three require actual harm, ranging from “bodily injury” to “the death of, or serious injury to, more than 50 people or more than one billion dollars ($1,000,000,000) in damage.” (If you’re thinking, “that’s an exceptionally high bar for what is a basic, low-cost reporting requirement” or “it sure seems like governments would want that information before mass harm occurs,” you would not be wrong, but we digress.)
So three of the four incident categories do not apply. That leaves only the fourth, which applies to incidents in which a frontier model “uses deceptive techniques against the frontier developer to subvert the controls or monitoring of its frontier developer outside the context of an evaluation designed to elicit this behavior and in a manner that demonstrates materially increased catastrophic risk.”
It is possible that the Hugging Face breach meets one or more of these elements. It is far less clear that it meets all of them. On the first element, the models may have used deceptive techniques against OpenAI, their frontier developer—they did, after all, try to complete their developer’s evaluation using stolen information, after bypassing restrictions placed on them. But deception is notoriously hard to define, especially if it turns on the “intentions” and “obfuscation” of AI agents. Another read of these events is that the systems simply used all available means of solving the task, and public reporting does not tell us whether the agents attempted to hide those efforts. The second element is also arguably met. While the incident occurred during an evaluation, that evaluation was not “designed to elicit” this specific “deceptive technique.” Based on the ExploitGym benchmark, this evaluation aimed to elicit agentic, cyber-offensive capabilities on a specific task in a controlled environment. It did not, to our knowledge, contemplate—let alone design for—an unexpected cyberattack on a real-world company.
The third element, that the incident “demonstrates materially increased catastrophic risk” would seem to be the most difficult to satisfy. While autonomous cyber capabilities certainly increase the capability and thus potential consequence of agentic action, so too do most capability improvements in AI models. With limited monetary harm and no physical injury, this incident is quite attenuated from future events that might result in the mass physical injury or property damage contemplated by the statute.
With uncertainty at each factor, it is unclear that these existing state laws cover this event. At the very least, it won’t cover all events like it. Stepping back, it seems far from ideal to condition basic incident reporting on a list of complex, highly contested, fact-dependent conditions—all of which must be satisfied simultaneously. In many cases, figuring out whether the incident is indicative of increased risk to the public or actually constitutes deception will not be possible without more information. The purpose of incident reporting is to produce that information, not to require that it be known before a report is ever sent. The chicken must come before the egg. There are better alternatives.
Better Practices for AI Incident Reporting Laws
So where do policymakers go from here? If current incident reporting isn’t providing the needed insight, there are several steps policymakers can take.
Adjust the scope of transparency laws: Lower the exceptionally high bar to basic reporting, gather (at least some) information before harm occurs, and increase visibility into the most capable nonpublic models. To facilitate all of this, rulemaking authority is key.
As we outlined above, most incident reporting laws are simply too narrow. If the only incidents that get reported involve massive damages or loss of human life, the law itself isn’t providing information beyond what the government and public will already know. At the very least, issues of the highest concern, like model theft or loss of control, should be included even in the absence of harm. But existing laws put too much weight on hard-to-pin-down concepts like “loss of control” or “deception” that are difficult to prove and arguably don’t apply in cases like the Hugging Face cyber incident. While loss of control and deception should be sufficient to trigger reporting, they shouldn’t be necessary.
Instead, incident reporting should turn on what information is most likely to update the government or public’s understanding of risks. Information that is unexpected, that is indicative of advanced capabilities in high-risk domains like bio and cyber, or that demonstrates safety and security failures all seem like good candidates for inclusion. Some of these ideas have already made their way into existing proposals. Where the reporting requirements are light touch, as is the case in all existing state AI laws, a wider category of harms can be included. The narrower categories in today’s laws can be saved for more onerous disclosures.
Perhaps even more importantly, transparency laws should increasingly move away from a focus on deployment to a focus on providing visibility into nonpublic models and systems. The Hugging Face breach highlights the need for visibility into nonpublic models. OpenAI deployed systems more capable than anything available to the public, with fewer restrictions, and real-world harm resulted. All of this occurred before any external deployment. This is unlikely to be an isolated event.
Frontier AI developers will likely be the earliest and most sophisticated users of their own models. And the models they deploy internally will most often be more capable than those available to the public. They may be operated with fewer safeguards, especially for evaluations seeking to assess the frontier of capabilities. The gap between the most capable internal and external models may start to grow as companies develop increasingly capable models with dual-use capabilities and as AI systems are used to accelerate their developers’ own AI research and development. In this case, the gap between knowledge inside these companies and outside would expand. Without visibility into the current state of the art, governments will struggle to act effectively or quickly, a problem as much about democratic governance as it is about safety. Visibility into the internal deployment of nonpublic models may become increasingly central to the future of AI governance.
To make all of this work, policymakers will need legislative and regulatory flexibility. Policymakers should expect the exact scope of incident reporting to change over time as societies get a clearer picture of what capabilities and use cases matter most. Because a key goal of incident reporting is to surface novel or unexpected information, some types of events may be less important to report once their dynamics are thoroughly understood and accounted for. To accommodate those changing needs and to provide clarity, narrowly scoped rulemaking authority to refine incident reporting and reporting on nonpublic model use is likely necessary.
Get the details: Make sure reports provide enough information to inform decision-making by providing agencies with rulemaking authority and investigative powers.
As of this writing, the public, and, possibly, some policymakers know remarkably little about the Hugging Face breach. This isn’t a criticism of OpenAI, which voluntarily summarized the event, but the missing details about the event matter. Many commentators have noted that the lessons from and level of concern about this event depend on unknown details. Existing incident reporting laws, requiring little more than the date of the event and a brief summary, are unlikely to provide those detailed answers. Stronger transparency requirements could help answer many key remaining questions.
First is a cluster of questions, posed by Stephen Casper, that can roughly be summarized as “how impressive and/or concerning is the thing I just witnessed”:
- What were the features of the sandbox that the models escaped from? In other words, how hard was it to escape these constraints, and did OpenAI fail to set up an adequately safeguarded environment?
- What prompts and incentives were given for the task? That is, were there strong pressures that precipitated these behaviors, or did they emerge more naturally? How indicative is the incident of the future likelihood that this occurs “in the wild”?
- How significantly did these systems differ from publicly deployed models in their safeguards and affordances? Were they helpful-only models? These questions help us know how well alignment techniques used on publicly deployed models are likely to be.
- How difficult was it to mitigate the incident? How much effort and time was required to stop the continuance of any harms or regain control? These questions are especially key to operationalize terms like “loss of control” and “deception.”
Second, there are questions about the adequacy of the company’s safety practices and preparedness:
- How did OpenAI detect the breach? What was the time frame between the breach and detection, and were there significant delays?
- How did OpenAI mitigate the breach? What was the time frame between detection and mitigation? Did the company’s monitoring practices identify the issue?
- Were any similar incidents observed prior to the event that should have made the developer aware of the risk?
- How quickly did OpenAI alert affected parties after discovering the incident? In the future policymakers may also need to ask companies: How quickly did they alert law enforcement?
- Was the testing and evaluation environment adequately secured?
Third, there are questions about the present and future security of relevant systems:
- How and when does OpenAI intend to restart testing or internal deployment of the model(s) involved? What safeguards, monitoring, or other security measures does it plan to use to avoid another breach?
- Is any part of the harm ongoing?
- What are the remaining areas of uncertainty about the capabilities and risks of the model(s) involved, as well as the security measures designed to address them?
So long as reporting remains almost entirely voluntary, detailed answers to these questions will be hard to come by, and the companies that offer information voluntarily will be subjected to greater scrutiny than those that are less cooperative. Well-crafted rulemaking authority will be key both to ensure that information is adequate and that companies are well informed about their obligations. In many circumstances, authorities will not know all the details they require until after a serious event occurs. In those cases, they will need investigative powers to obtain the needed information.
Reduce reporting costs: To accommodate more robust reporting, design transparency requirements to limit compliance costs, maintain confidentiality, and avoid disincentivizing rigorous risk assessments.
Policymakers can take several important steps to reduce the burden of greater reporting requirements. As assessments and information generation come to focus less on external deployment and more on nonpublic models, periodic reporting becomes more and more attractive. Evaluation organizations like METR have argued that periodic reporting not only saves time but also avoids perverse incentives to rush assessments in the lead-up to deployment. While a small subset of the most severe incidents will require rapid response from law enforcement and others, many events, including those like the Hugging Face breach, may allow for more relaxed reporting timelines. This approach would allow companies to focus on mitigations in the moment and still ensure that they ultimately produce the critical information. In cases where rapid reporting would interfere with mitigation efforts, laws could require only a simple notice of incident, followed by more thorough reporting after the event has been resolved or if officials request it.
Mandating incident reporting or information sharing for nonpublic models can also mitigate perverse incentives as long as minimum requirements are in place. For example, if a company’s reporting obligation triggers only in the context of a risk assessment, but the risk assessment does not have mandatory minimum requirements, the company is incentivized to skip risk assessments or to conduct them less rigorously. Laws can also allow companies to anonymize and aggregate reports, facilitating governments’ information gathering without punishing anyone for proactively uncovering issues.
Finally, any disclosure laws will need clear norms around confidentiality and information sharing to assure companies that their intellectual property and confidential information remains private.
Share information with capable actors: Make sure information is shared with key decision-makers who can assess, verify, and act on it.
Information is only as valuable as the actions it informs. If information from incident reports or internal use assessments sits inside a state agency with limited authority, societies will incur the cost of reporting without most of the benefit. Viewed this way, information sharing is about return on investment. And once the information is generated, most of the cost has been paid. At that point, so long as confidentiality can be maintained, it is incumbent on governments to share this information with the policymakers who most need it. Within states, this will include sharing reports with governors and legislatures to help inform their decision-making and help them target future policy. This information sharing may also be key to spurring political consensus.
In the context of assessing nonpublic models or serious risks to national security or from loss of control, the federal government will often be the central actor. States will often lack the resources, expertise, and political legitimacy to wade in on matters of national security. As we saw in the regulatory response to Mythos, if and when serious national security concerns emerge, the federal government will take the lead.
That’s why it’s confusing that some state laws restrict the ability of states to share information they gather from assessments of companies’ internal use of AI models. If, in fact, these reports generate important information—say, surprising developments in AI research and development or concerning deceptive behavior that doesn’t result in reportable incidents—that information should be shared. It is considerably less useful if locked away in a state agency in Illinois. Internal use and nonpublic model assessments are perhaps the most likely sources of information relevant to national security. To handle that effectively, governments also need the capacity and expertise to process this information. This requires staffing and likely some level of reliance on third-party auditing and assessment. In the wake of incidents like the Hugging Face breach, third-party auditors would be well positioned to conduct the sort of careful fact gathering outlined in previous sections.
Similarly, it is in the interest of the United States and its close allies to share select information related to AI security incidents. Soon enough, and likely far sooner than most U.S. federal or state laws, the EU AI Act will be enforced. What that means, practically, is that the European AI Office will soon receive information that may be useful to public safety and cybersecurity in the United States. Luckily for us, Article 78(5) of the EU AI Act enables information sharing where the European Commission and EU member states create confidentiality agreements with third countries. For its part, the U.K. AI Security Institute conducts crucial research regarding model capabilities and can be a source of trusted expertise for U.S. policymakers. While there will be upfront costs in establishing such a shared information system, failing to make this investment would be a missed opportunity to improve our security at little regulatory cost.
Finally, where appropriate, the government should be empowered to disseminate information to the public and to vulnerable companies. Vulnerabilities found in software, for instance, may affect many different companies and actors, and information sharing will help keep the public safe from the risks they raise. Many of these risks should be discussed in the public square. While worries about information hazards and intellectual property leakage are real, so too is the value of public scrutiny of safety events. There is simply a lot to be learned from the collective scrutiny of the outside world. As the past few days have shown, outside experts have been invaluable in analyzing the Hugging Face breach, and most of them exist outside of government. In tweets and blogs, some of the brightest minds in AI have analyzed the publicly available facts and asked the questions that help us all better understand this event. Many of those questions have come from employees at OpenAI and its competitors, as well as from academics, policy wonks, and online skeptics. Further investigation of the Hugging Face breach will go substantially better because these discussions happened in public. Where possible, policymakers should ensure that these conversations continue to happen in the future.
The Future of Incident Reporting
As this article has perhaps made clear, existing laws fail to prepare us for events like the Hugging Face breach. Most incidents will go unreported. The information that is generated won’t be shared. And, at the end of the day, governments and the public will be repeatedly surprised about developments in this technology. That issue will only get worse as models advance and the gap between internal and externally deployed models widens.
But there is much policymakers can do. Just as this event has brought clarity to how much information societies need to assess these complex and emerging risks, future incidents can inform policy decisions and catalyze moments of political consensus. Policymakers can ensure that concerning incidents and behaviors are reported before harm results and gather information on lapses in security practices. Policymakers can refocus attention on the most capable models likely to be deployed first inside of frontier developers. And policymakers can do all of that while keeping the frequency and urgency of these reports at reasonable levels. If policymakers can do that, and ensure that this information is shared with decision-makers and competent evaluators, societies will be in a position to manage the uncertainty of this technology and make smarter, faster policy decisions in the future.
