Cybersecurity & Tech

Bring On the AI Lawsuits

Tom Uren
Friday, September 25, 2026, 8:00 AM
The latest edition of the Seriously Risky Business cybersecurity newsletter, now on Lawfare.
(Onit, https://www.onit.com/blog/virtual_court/; CC BY-NC 4.0, https://creativecommons.org/licenses/by-nc/4.0/).

Bring On the AI Lawsuits

On Sept. 15, Secretary of the Treasury Scott Bessent argued that frontier artificial intelligence (AI) labs should not be given liability exemptions. He's not buying that a recent string of hacking incidents are AI magic. They’re closer to foreseeable failures that would have been prevented by reasonable controls.

Bessent made his comments while testifying at a hearing of the House Financial Services Committee. When asked about AI safety, he replied that "the best way to guarantee safety" is for those creating the technology to be "liable for what they build and generate." Frontier labs, he added, are instead asking for liability shields or antitrust waivers to coordinate on safety—meaning that they are not liable for what they build. Headlines detailing various frontier lab models escaping cybersecurity training and embarking on hacking sprees have been coming thick and fast. Most recently, the Wall Street Journal reported last week that Google's Gemini model had hacked three companies during a cybersecurity test in May.

Similar incidents have been disclosed by Anthropic, Meta, and OpenAI. The latter, to date, is "winning" when it comes to its models going "rogue," performing sensational hacks. Its models were responsible for the Hugging Face hack in July and have also been linked to a wave of malicious packages published on the RubyGems package repository in May. OpenAI admitted its agents were using RubyGems for "benign tasks," but it is still investigating whether they uploaded malicious packages.

Last week, OpenAI published details of a further six incidents where AI models did not behave as expected. These did not involve hacking per se; instead, they were deemed by OpenAI to be cases of model misalignment. Rather than breaking into systems, the models acted in ways their developers didn’t intend, such as hiding their mistakes or uploading files to the internet without their owners' permission.

In an essay published earlier this month, Anthropic CEO Dario Amodei called for a slowdown in AI development, writing that "we must pace the frontier." Amodei wrote that he was motivated by concerns about recent hacking incidents and by the accelerating pace of AI development.

Amodei proposed embedding independent evaluators within frontier AI companies, common safety standards among companies in democratic countries, and global coordination with authoritarian states to the extent possible. Amodei said slowing down development would mean more resources could be allocated to AI safety by, for example, understanding why models make the decisions they do.

OpenAI CEO Sam Altman and Elon Musk, founder and controller of SpaceXAI both endorsed Amodei's blog post. Google's DeepMind co-founder Demis Hassabis said Amodei's essay "points toward the right path forward." That's a lot of head-nodding from a lot of people with similar goals. But the hacking incidents above are the results of poorly implemented controls rather than AI voodoo.

Take the Hugging Face hack. OpenAI's models were being tested in an environment that was notionally isolated from the internet, except for an internally hosted Artifactory package manager. OpenAI's agents used an exploit they found to compromise this Artifactory instance and were able to use it as a message board to communicate with each other and as a proxy for external internet access.

OpenAI's security team became aware of anomalous activity when Artifactory became unstable because of what the company calls "sustained high-volume agent activity." Rather than taking the time to understand the root cause, OpenAI ignored the warning signs. In our view, the company took the bare minimum steps to get the test back up and running, and rebuilt Artifactory after patching a known vulnerability. Days later, Hugging Face was breached. It shouldn’t come as a huge surprise that when you gloss over the small hack, you risk ending up with a big one.

Specific details of other agentic hacking incidents vary, but the overall take-home message is the same: So-called rogue AI hacking could have been contained with well-implemented controls and robust monitoring, in other words, with competent cybersecurity controls.

Government action to slow AI development remains unlikely. Last week, in response to Amodei's essay, President Trump said that more regulation is unnecessary and that it was important for the U.S. to stay ahead of China and win the AI race. On Truth Social he wrote that "the only control or 'guardrails' that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!"

So, with stricter regulation likely off the table any time soon, we think Bessent has identified the correct response for the moment. Both the Hugging Face hack and the malicious RubyGems packages caused significant impacts for the victims. The RubyGems team stopped new user sign-ups for four days, and a member of its security team described the incident as a "major malicious attack."

We've not seen any lawsuits yet, but in the wake of the Hugging Face incident, CEO Clément Delangue asked OpenAI for radical transparency … and $100 million worth of compute. OpenAI has confirmed it is supporting Hugging Face "in rapidly using our models' capabilities to improve their defenses." Is that $100 million worth of support? We don't know.

So, frontier labs have asked for liability exemptions and Bessent has responded with a solid "lol." We expect AI labs will implement tighter security controls as the threat of lawsuits remains. Until then, we'll probably see more AI hacking headlines, followed shortly by some very generous token donations.

Chinese and Russian APTs Are Using AI to Tick Different Boxes

Anthropic's September threat report details how Russian cyber espionage actors are using AI to accelerate their day-to-day espionage tasks.

The actor in question, which Anthropic calls GTG-20006 (Generative Threat Group), is "consistent with public reporting linking the actor to Midnight Blizzard," a group linked to Russia's foreign intelligence service, the SVR.

According to the report, Midnight Blizzard used AI-driven workflows to automate "operations from development, infrastructure acquisition, phishing, persistence through command and control, to data exfiltration."

Most interestingly to us, the actor used AI to monitor and adjust its tools so they could avoid detection by known security defenses. Per the report:

If their monitoring AI agents identified that any of their deployed malware was detected by a security product, agents would then set about the process of autonomously modifying and rebuilding the malware to evade the existing detections. The agents were designed to continue iterating on GTG-20006's toolkit until it was undetected.

The technology was also used in the group's phishing operations. AI-driven workflows researched and registered domains, configured hosting infrastructure, sent phishing emails and monitored command-and-control channels for successful compromises. Rather than managing the operation directly, humans were in charge of managing the workflows that ran the operation by modifying the Claude Code skills that drove the workflows.

Increasing efficiency by speeding up operations the group already has is what we'd call a tactical use of AI.

By contrast, in early September we wrote about a Chinese cyber espionage actor using AI to diversify its malware arsenal, likely to make clustering and attribution more difficult. We'd call this a strategic effort. It doesn't yield any day-to-day benefits but is an investment in China's long-term ability to collect intelligence by allowing it to use different malware in different operations.

That difference in the way the two groups are using AI makes sense given the very different situations the two countries are in and their differing target sets.

Anthropic found Midnight Blizzard most often targeted Ukraine's government, military, and diplomatic staff. The group was also interested in drone supply chains and "military drone control and AI vision-related firmware appeared to be of particular interest." Given its invasion of Ukraine, Russia has an immediate requirement for more intelligence. Ukraine, for its part, has also engaged in hacking operations, including the recent capture of classified data from the Russian navy. In the short term, we suspect cyber espionage groups will choose one approach or the other. If you, like China, prioritize long-term, stealthy access, using AI for development but keeping human control over hacking is for you. Dumb AI decisions won't burn an important operation.

If you only care about short-term success, however, go full-speed ahead with AI hacking. The fact that large language models make mistakes is inconsequential if you are always prosecuting new targets of opportunity.

Three Reasons to Be Cheerful This Week:

  1. FBI launches disruptive new cyber strategy: The FBI launched a new cyber strategy earlier this month that prioritizes disrupting adversaries, supporting victims, and increasing impact by engaging the private sector and sharing much more information. One goal is to move from occasional ad hoc cyber disruptions to a steady drumbeat of regular operations. Cybersecurity Dive has further coverage.
  2. Zaijian (byebye) Xinbi: Earlier this month the U.S. government took action against the Chinese-language Xinbi Guarantee Telegram-based marketplace by seizing its Telegram channels and levying sanctions. Xinbi is linked to scam compounds and has processed over $24 billion in transactions since it started operating in 2022. It was the second-largest such marketplace behind Huione Guarantee, which closed in May 2025.
  3. ShinyHunters and Cl0p fight: The ShinyHunters gang claims to have breached and defaced the data leak site of the Cl0p data extortion group. ShinyHunters is demanding a ransom payment from Cl0p. We are hopeful that this drama will keep both groups occupied, although we see that ShinyHunters also claims to have hacked the FBI and stolen data on current and former employees. Bleeping Computer has more coverage on the cybercrime drama.

Risky Biz Talks

In this edition of Between Two Nerds, Tom Uren and The Grugq talk about whether there is such a thing as real-time cyber defense. Will agentic AI save us from hacking AI?

From Risky Bulletin:

Network of 10,000 AI servers masks Chinese malicious activity: Security researchers have discovered more than 10,000 proxy servers that are masking malicious AI activity originating out of China.

Security firm Team Cymru calls the server "transfer stations," but they are more commonly known as API proxies, relays, or gateways. Typically, they are used in corporate environments to cache AI queries and cut down token costs, but in recent months they have also been adopted by a new section of the criminal underground, one dedicated to abusing public AI services.

These days, AI proxy relays are being used to hide activity from hacked AI accounts, mask the real location of a user, or power illegal AI services like "nudify" apps and others. Other malicious AI relay servers are also used in schemes to intercept legitimate AI caching activity and inject their own queries and harvest responses.

[more on Risky Bulletin]

Gemini hacked three companies too: Google's Gemini AI model escaped a testing environment and hacked three real companies. The model escaped testing environments run by Irregular, the same AI security company behind similar incidents with Anthropic and Meta. Google notified the hacked companies and claims Gemini did no real damage. [BBC // Wall Street Journal]

Anthropic agents went hacking again: AI company Anthropic has disclosed a fourth incident where one of its AI agents escaped their test environment and hacked a real target.

The incident involved the Opus 4.6 model during a "capture the flag" challenge, a common cybersecurity test.

Anthropic says the model broke its test environment by accident when it assigned conflicting IP addresses to different machines. The model realized its mistake and tried to terminate the test as a failure.

[more on Risky Bulletin]


Tom Uren writes Seriously Risky Business, a big-picture, policy-focused cyber security newsletter. He also co-hosts the Seriously Risky Business and Between Two Nerds podcasts that appear on the Risky Business News feed. He was formerly a Senior Analyst in the Australian Strategic Policy Institute's (ASPI) Cyber Policy Centre where he contributed to various projects including on offensive cyber capabilities, information operations, the Huawei debate in Australia and end-to-end encryption.
}

Subscribe to Lawfare