The AI Model Distillation Paradox
On Sept. 1, the Department of Justice filed its brief for the artificial intelligence (AI) companies’ right to appropriate and reuse your work—which it called a “statement of interest”—in OpenAI’s copyright battle against the New York Times and other publishers.
The federal government’s first direct intervention in the wave of lawsuits over generative AI training data took the position that feeding copyrighted text into an AI model for training purposes is fair use. OpenAI's training is “exceedingly transformative,” it argued, and requiring companies to license everything they ingest would “render [the] training of AI models impermissible.”
At stake, the department warned, is all America holds dear: its freedom of expression, its lead in technological innovation, even its national security. By handing a “competitive advantage to foreign adversaries who are not so encumbered,” the department contended, the law would render unto China what should be America’s.
Eight days later, however, the National Security Agency, the Cybersecurity and Infrastructure Security Agency, and the FBI were singing a different tune. In a joint cybersecurity advisory report, the government accused six Chinese laboratories of running “industrial-scale distillation campaigns” against American frontier models. Distillation is a method that trains a smaller AI model to mimic the behavior of a larger, more powerful one, allowing it to run on cheaper hardware without sacrificing performance.
In its brief, the government alleged that the malefactors—Chinese firms such as DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.ai—extracted billions of tokens, the basic data units that an AI model reads and generates, from variants of such wholesome American products as Claude, GPT, Gemini, and Grok, “likely with Chinese government awareness.”
This was no occasional shortcut: According to the cybersecurity advisory, “aggressive, malicious, and targeted distillation activities” began at least as early as 2024, and formed the “core” of Beijing’s AI strategy.
When American labs ingest the world’s writing without permission, it’s hailed as transformative innovation. When Chinese labs do the exact same thing to U.S. models, it’s a “malicious and targeted” campaign to “to extract proprietary functionality and reasoning capabilities”.
All of which raises three critical questions:
First, what actually separates training a model from distilling one?
Second, given that both practices rely on taking someone else’s work without consent, why does the federal government regard American and Chinese companies so differently, and is there any legal or rational basis for that difference?
And third, if there isn’t one, how should the law treat the mass harvesting of data—whether the harvested material is written by humans or generated by machines?
Training vs. Distillation
An AI model processes an input—a simple question like, “What color will the sky be tomorrow?”—into a corresponding output using a vast architecture of numerical parameters called weights.
To visualize this process, imagine a giant mixing board in a recording studio. The user’s prompt is the raw audio flowing in—the vocals, bass, or drums. The weights are the thousands of sliders and knobs speckled across the console—some dialing up the bass, and others muting the treble or adding effects. The final song reverberating from the speaker is the model’s answer.
No human sound engineer adjusts the dials by hand. A large language model simply has too many parameters—literally billions, sometimes trillions of them.
Instead, the AI labs set the weights through a process called training. The developers initially set them at random, and the console thus produces little more than static noise. But they then expose the model to oceans of human data—raw text and code, and, for some models, images, video, and audio. And they ask the model, over and over again, to predict the next “token,” for instance, the next word in the sentence, “The sky will be…”
After each guess, the engineers compare the output with the “right” answer (“blue” in pleasant weather, or “gray” in less fortunate climates).
With every misstep, billions of invisible knobs turn collectively toward an arrangement more likely to produce the correct answer. Repeat this process enough times and over vast enough oceans of data, and the static noise slowly transforms into music.
In a later phase, known as fine-tuning, engineers reward desired answers for helpfulness, clarity, and safety, and discourage discordant ones, thereby shaping the model’s behavior and aligning its “voice” with human ethics and values—so that your ChatGPT instance doesn’t sound sociopathic when you ask it a question.
But now imagine a lab that wants to train a capable model quickly and cheaply. It doesn’t have the time or resources to ingest so much data or tune so many parameters, and besides, other labs have already done that work. Is there a way to harness the work other companies have already done and not recreate the wheel? It turns out that, just as OpenAI and Anthropic can teach their models using all of recorded human knowledge, DeepSeek and Z.ai can train theirs by using ChatGPT and Gemini.
This is the process known as “distillation.” Geoffrey Hinton, the so-called Godfather of AI, explains the technique as an agile “student” model trained to imitate the behavior of a complex “teacher” model. Rather than relearning the universe from scratch, the student repeatedly listens to the master’s performance, learning to mirror its responses. This approach produces a leaner and faster model that captures much of the original genius, while dramatically reducing the latency and hosting costs. It also bypasses the need for extensive, human-annotated datasets.
And if that sounds familiar, that’s because at its core, distillation is simply a specialized form of training. Both processes take an existing fountain of written and recorded work and pass it through a transformative technical crucible, thereby teaching a new model to shadow the aggregated brilliance of what it has just heard.
The most fundamental difference lies in the source of the fountain appropriated for the training. In distillation, the student model learns from another machine, rather than directly from the humans whose work that machine had previously absorbed.
A Legal Gray Zone
Lest you think that distillation is the disreputable Chinese stepchild of wholesome American training, U.S. frontier labs engage in distillation too—every day. Far from an exotic practice, distillation is part of the bread and butter of all AI. Every major American lab distills its own models to train lighter, faster, and more energy-efficient generations. Stanford’s Alpaca and Microsoft’s Orca trained smaller models on the outputs of larger ones, and the creators shared their methods openly with the research community. Under oath in April, Elon Musk conceded that xAI had trained Grok on OpenAI output, calling it “a general practice among AI companies.”
Yet, when Chinese labs run the same process, the whole thing surreptitiously transforms into capability theft. The government’s cybersecurity advisory describes an elaborate web of evasion, with distillation requests scattered across platforms and providers such that no single operator can recognize what is happening. The companies allegedly aim their extraction at the chain-of-thought reasoning that labs deliberately conceal, as though that were not precisely what domestic labs are doing with literature and journalism. Remote cloud providers and third-party aggregators allegedly strip user metadata from output to disguise it. And a gray market of proxies—known as “transfer stations”— bypasses geographic restrictions and avoids traceability, as though the American labs had been transparent with copyright holders about the origins of the material they were using without consent.
All of which makes distillation sound very nefarious, while not actually answering the key question: whether learning from an unconsenting model’s outputs, as opposed to from an unconsenting human’s output, is a form of theft.
On that question, the law is ambiguous. Reverse engineering a model, generally speaking, is lawful, and copyright rarely reaches AI-generated outputs, as copyright is reserved for human creation.
Legal scholars have debated whether distillation violates the Defend Trade Secrets Act, which allows companies to sue over theft of trade secrets. But claiming trade secret protection for text sold with an API key is a tricky business, because commercializing a secret almost always destroys its legal protection. U.S. courts, for their part, remain undecided on the issue. Referring to distillation as “stealing” is thus more of an ideological statement than a factual description of the law, which regards it, rather, as just another form of learning.
What distillation clearly violates, however, is the labs’ terms of service. OpenAI, Meta, and Anthropic have written prohibitions on the practice directly into their user agreements. These provisions explicitly prohibit the use of platform outputs, APIs, or services to train rival models through any kind of distillation or reverse-engineering. So there’s little doubt that the Chinese companies, to the extent the U.S. government allegations are true, are potentially liable for a contract-breach suit for their conduct.
Yet despite having built that private legal regime directly into their sale of service to the Chinese providers, the American companies have been reluctant to enforce it—even after claiming to detect massive distillation campaigns.
So why aren’t the American labs suing the Chinese labs for distillation?
In January 2025, OpenAI and Microsoft concluded that a DeepSeek-linked group had siphoned data through their API. The companies accused the Chinese company of breaching terms of service but did not sue. In September 2026, Anthropic traced nearly 200 million exchanges to distillation campaigns, including an Alibaba-linked operation peaking at 3 million queries a day across more than 3,500 fraudulent accounts, and a Moonshot-linked effort that harvested live customer conversations through Claude, sometimes exposing sensitive data or credentials, in order to train its own Kimi model. Anthropic published a threat report, but it did not sue.
These are not companies shy about asserting their legal rights. But suing over distillation carries risks for them. They would, if they went to court, be asking the courts to issue rules to which they themselves would be bound. If those rules were written broadly enough, they could theoretically limit practices that the labs themselves routinely engage in r. A ruling, for example, that terms of service can govern machine outputs would effectively amount to an intellectual property right in model outputs. That would be a verdict that the publishers currently suing American labs for the “largest theft of labor in history” would love to carry into the next court hearing—one in which the American labs would have to argue that such theft is just fine.
Glass Houses
These are, indeed, treacherous waters to go wading in. The models that the federal government now defends as strategic assets were themselves built on the texts, images, sounds, videos, and code of millions of creators who were neither asked nor paid. OpenAI transcribed more than a million hours of YouTube videos to train GPT-4, in plain breach of YouTube’s terms, which ban automated harvesting. Google, which owns YouTube, was reported to have performed the same practice on its own platform, and allegedly broadened its terms of service so that public documents on Google Docs and reviews on Google Maps could be used to train its Gemini family of models.
Anthropic went a step further. Under a confidential program referred to as Project Panama, the company bought and destroyed millions of antique and secondhand books to train its models. Vendors reportedly sheared off book bindings using hydraulic cutters before feeding loose, century-old pages into high-speed industrial scanners. Data, like gold and oil, is a finite resource, and for a frontier lab like Anthropic, 18th century literature is the equivalent to striking a gold mine. Ingesting these rich, historical texts is essential to maintaining the quality of Claude and preventing the decay that happens when an AI model is nourished on synthetic data and the internet’s exhaust.
Before that, the company had simply downloaded millions of e-books from shadow libraries such as LibGen, which provide free access to digital books and academic journals that would usually be locked behind paywalls and restricted by copyright.
In June, Anthropic wrote to the Senate Banking Committee to report that Alibaba had run 28.8 million Claude queries across 25,000 fraudulent accounts, constituting “the largest distillation attack on Anthropic to date.” Four weeks later, however, a court approved the $1.5 billion settlement the lab would pay the authors whose books it had used without permission.
The company complained to Congress both in its capacity as a victim and in its capacity as defendant—both about what had been taken from it by Alibaba, and about being forced to pay for what it had taken from countless others—all in a single summer.
So, What’s the Excuse?
There is one big difference between distillation and training, and it’s not necessarily a flattering one to the American labs: The Chinese AI companies are actually paying for the service whose terms they are then violating. They are using real subscriptions to Claude and ChatGPT to train and scale their models based on data they purchase. The American companies, by contrast, have largely followed a policy of copy now, and figure out the rights later.
An additional factor further complicates the picture: Distilled model capacity does not stay inside Chinese borders. Rather, companies like Alibaba distill American models at marginal cost and then release the results as open-weight models, which are free for anyone to run and modify. The models are fast, flexible, and cheap—precisely the qualities that American developers want for post-training, the step that turns a raw, next-token predictor into an obliging, instruction-following assistant. And ironically, American companies then use these Chinese open-weight models to power their own post-training.
The result is a boomerang effect in which capability leaves the American frontier labs through an API, returns as Chinese open-weight infrastructure, and is then rebuilt into American products.
This helps to explain why the Chinese lab Qwen’s share of open-model fine-tunes—whereby an already-trained AI model is trained on a smaller dataset to adapt it to a specialized task—reportedly rose from 1 percent in January 2024 to roughly 69 percent by February 2026. A company in Silicon Valley post-training its model on Qwen is doing, one-step downstream, exactly what the federal government has identified as a national security threat when Qwen does it to Anthropic.
How, then, do American labs distinguish between what they do to train their models, and what Chinese firms are doing to distill them?
One argument made by frontier labs like OpenAI is that the data they use to train their models is transformative of the original material, whereas distillation copies a specific product’s behavior to build a competitor in the same market. Training on the world’s writing, OpenAI has argued, is “transformative” because ChatGPT is not a “replacement” for the New York Times. DeepSeek, by contrast, harvested ChatGPT’s architectural information to build a rival model. In fair use terms, this means harm to the market: An unauthorized use acts as a substitute for the original copyrighted work, reducing its current or future commercial value.
The problem with this line of reasoning is that the New York Times does argue that ChatGPT is replacing its journalism, though perhaps not the publication as a whole. In its copyright infringement lawsuit against OpenAI and Microsoft, the newspaper states that training models on millions of its articles is “substitutional” rather than “transformative,” since it allows readers to access near-exact copies of its reporting without a paywall. Fair use, moreover, applies only if there’s copyright to begin with, and model outputs largely fall outside of the world of the copyrightable to begin with. Legally, the Times argues, American AI companies thus cannot rely on the fair use doctrine to protect them.
The American labs also argue that their models are trained on public data on the open web, whereas Chinese labs deliberately manipulate access controls and gated APIs to get their materials. OpenAI’s memo to the House China Committee accuses DeepSeek of creating fake accounts and using disguised routers and third-party resellers. This was meant to evade the commitments that the Chinese companies had made when they signed the company’s terms of service, which largely do not apply to websites that are scraped from the internet.
When the DeepSeek engineers clicked “I agree” to terms banning training on outputs, they agreed not to replicate the OpenAI product. By contrast, in the wake of hiQ v. LinkedIn, American courts have been reluctant to treat scraping public pages as unlawful access if the scraper did not promise to play nice.
But this logic has vulnerabilities too. YouTube has terms banning harvesting, and LibGen was built not on public but pirated data. Ironically, Project Panama represents the cleanest case of data harvesting, since buying and destroying a physical book is a lawful, if morally dubious, activity.
Then there is the claim that Chinese labs are “free-riding” on American research and development. Rather than training models from scratch, companies like DeepSeek slash months off their development timelines by lifting capability directly from American models via API harvesting. What is being piggy-backed is not raw text or code, as is the case in training, but billions of dollars worth of compute and reinforcement learning. This is true. It’s also unambiguously the same claim that American publishers are making when they refer to “the largest theft of labor in history.”
The American labs further contend that distilled models lose the safeguards the American labs had meticulously built to prevent bad things from happening. In February, Anthropic argued that distilled models lose the protections it has built against its models being used to create bioweapons and for cyberattacks. It also warned that distilled models can subsequently be ingested into military, intelligence, and surveillance systems that it will not allow Claude to be used for.
These assertions may well be correct, and Anthropic’s warnings may indicate real dangers. But it is not an argument about the law. It is, rather, an argument about what happens to the model downstream. It would therefore apply as much to a model that Chinese labs built from scratch and simply didn’t build safeguards into.
Finally, there is the hawkish claim that China is a foreign adversary and that model distillation undermines U.S. export controls and national security. Anthropic has warned lawmakers that industrial-scale distillation attacks mask the true origin of Chinese AI progress, and that observers will mistakenly conclude that semiconductor export controls are failing as a result of Beijing’s domestic innovation.
Again, this may be correct. But if Chinese labs can reach near-frontier performance by distilling U.S. models, perhaps the export controls aren’t all they are cracked up to be. More than anything, that suggests lousy policymaking.
You Can’t Have Your Cake and Eat It Too
China, for its part, has taken a stricter line on data governance and copyright infringement than the United States. Beijing’s 2023 rules for generative AI require training data to come from strictly “lawful sources” and forbid model providers from infringing anyone else’s intellectual property.
Chinese courts have held an AI company liable for generating images of the Japanese superhero franchise Ultraman, and a Beijing court ruled that training a voice model on an actor’s recordings without consent amounts to breaking the law. A landmark ruling also recognized that an AI-generated image can be protected by copyright, and that a human prompter can be considered an author under law. American courts, by contrast, lean overwhelmingly on the fair use doctrine, with “transformative” being the magic word that renders ingesting the world’s writing lawful.
China appears to be stricter about what goes into a model and more lenient about what is taken out of someone else’s. The United States offers the reverse picture.
Last month, Washington insisted that training models on other people's work without permission is so obviously fair that licensing requirements would be ruinous to national innovation. It also insisted that training models on other people’s work without permission is an act of industrial predation posing a grave threat to national security.
Both propositions cannot hold simultaneously. Either the outputs of these systems are protectable under the law, in which case the writers, musicians, artists, and programmers whose work American models were built on can hold a claim that the Justice Department has denied. Or they are not, in which case the six Chinese laboratories accused of “malicious and targeted” distillation attacks have not done anything that Elon Musk has not freely admitted to under oath, without consequence.
Notably, the federal agencies have proposed neither a statute nor a rulemaking in order to prevent future distillation “campaigns.” The advisory's most concrete suggestion is for American frontier labs to “subtly alter” or poison their outputs—just enough to degrade the quality of the training data that Chinese companies are distilling, instead of blocking a suspected rival’s account, which would alert them that they have been caught. The government has refused to confront the most glaring challenge: what model outputs should be classified as under the law, and under which circumstances they could concretely be defended.
The distillation paradox thus extends beyond the American labs’ thinly veiled hypocrisy when it comes to training their models on the world’s work. It reveals a government unable to decide whether to restrict this practice in the name of national security, or encourage it in the name of transformative innovation.
