America Must Protect Its Training Data
Washington has spent the summer alarmed by Chinese artificial intelligence (AI). After the release of Kimi K3—a Chinese model that pulled within reach of the U.S. frontier models—the White House began weighing plans to ban enterprise use of Chinese models, and Congress has escalated probes into the use of Chinese AI by U.S. companies.
Kimi, like all AI models, was trained using two key resources: semiconductor chips and training data. Policymakers have focused heavily on limiting China’s access to the U.S.’s best chips— producers such as Nvidia and Intel face myriad export controls restricting sales to China.
Policymakers, by contrast, completely ignore data providers. These are U.S. companies—including Scale AI, Surge AI, and Mercor—that work with AI labs directly to build specialized training data. Gone are the days when AI companies could collect enough data by simply scraping the internet. Now, an ecosystem of firms, each worth tens of billions, builds simulated workplaces where AI agents practice real jobs, recruits thousands of specialists to produce expert-written demonstrations, and sells customized training data for billions every year.
Despite all this, virtually no laws govern data providers’ dealings with Chinese AI companies.
And China’s labs have taken advantage. Recent reporting puts the top six Chinese labs’ annual spending with U.S. AI data companies at over $500 million. And experts have long speculated as much: SemiAnalysis wrote in January that Surge sells training environments “to Chinese labs like Moonshot and Z.ai,” and that access “played a huge part in increasing the capabilities for Kimi K2 Thinking and GLM-4.6.”
Given their importance, data providers should be restricted from selling to Chinese AI companies. High-quality data is an essential ingredient of AI training, and it’s the scarcest. Creating it demands increasing sophistication and compute, which is why the U.S. data pipeline is the best in the world. These American firms should not be subsidizing Chinese AI progress, and certainly not on the backs of U.S. talent.
Fortunately, policymakers can close this gap—and successfully enforce it—with tools they already have. The White House should extend the Department of Justice’s existing bulk data restrictions—which already bar the sale of Americans’ personal data to China—to cover AI training data as well.
Data Is a National Asset
Understanding why training data is a strategic asset requires understanding how it has changed. Early models learned by digesting raw internet text—such as articles, forum posts, and even Wikipedia entries. But in an age of AI agents, the frontier has shifted from static text to purpose-built environments.
Today, when an AI lab wants its model to get better at something—medical diagnosis, for example—it hires a data provider to build specialized training environments where the AI practices and learns how to perform a suite of tasks.
Every part of these environments must be manufactured: the software where the agent acts, the task suite, and the verifiers that determine success conditions on each task. And subject matter experts have to be involved in this manufacturing process as well. For example, to build one recent evaluation suite, Mercor assembled 256 professionals who averaged nearly 13 years of experience—former consultants from McKinsey and BCG, bankers from Morgan Stanley and Citigroup, corporate lawyers from Disney—and had them spend five to 10 days constructing simulated firms, complete with internal files and working software tools, in which AI agents could be trained and tested on real professional work.
These expert-curated environments are an important driver of progress for AI agents, and AI companies are willing to pay for environments that simulate any important modality. According to Epoch AI, a high-fidelity clone of a Slack-style workplace costs around $300,000, and labs spend billions every year for the construction of more environments.
And the labs don’t order just once. They run experiments, find what the model still can’t do, and come back asking for more data to fill the gaps in the environment: new simulated workplaces, better ways of grading practice attempts, and new pipelines to human experts.
Years of these feedback loops have concentrated institutional know-how in a few key vendors. By one industry count, more than 75 percent of this market belongs to just four U.S. companies. It’s simply too costly to switch away from the firms with this know-how: Epoch AI, after interviewing 18 practitioners across the industry, concluded that “maintaining quality while scaling is the number one bottleneck …. Finding the experts isn’t that hard, but managing them and doing quality control is hard.” A former Google reinforcement-learning researcher, Auriel Wright, was blunter about what happens when a vendor gets it wrong: “Researchers don’t want your broken RL environments because they will make our models worse.” Buying low-quality data means the model learns the wrong things and the training run has to be thrown away.
The market also reveals how much of an advantage it is to have access to high-quality data that others don’t: Environments sold exclusively command four to five times the price of those sold nonexclusively. Unlike the pretraining regime, the current regime of AI training data is one where not everyone has access to the same corpus.
Put succinctly, the training data powering today’s AI agents is the product of a few sophisticated firms whose product: (1) warrants billions in annual spending; (2) cannot be easily replaced by others; (3) companies pay a premium to exclude others from; and (4) has been carefully designed alongside U.S. professionals. This is what the White House’s AI Action Plan means when it calls high-quality training data “a national strategic asset.”
It is also worth noting that building high-quality training data also requires compute, not just sophistication. Vendors now use frontier models to generate data and judge other models’ attempts—Mercor’s chief executive recently confirmed the company spends more on tokens than on employee head count.
Indeed, the institutional knowledge and compute required to create AI training data is why China’s own data ecosystem remains underdeveloped: A Z.ai cofounder complained in People’s Daily that China’s high-quality data is “fragmented and scattered,” and the Center for a New American Security finds that China’s data industry is immature enough that its labs have to resort to harvesting U.S. model outputs, which “spares Chinese developers’ own limited compute for other uses,” allowing China to be an “even faster follower.” Nathan Lambert, an AI researcher at the Allen Institute who visited most of the leading Chinese labs this spring, came back struck by how there was “almost no data industry” comparable to the United States’.
That’s why Beijing should not have access to the AI training data we build. The institutional knowledge these firms have accumulated to build training data is a strategic asset. And even where the data itself could be replicated, China would have to burn its own scarce compute and labor to do so.
China itself has begun to see just how strategic an asset AI training data is: Its commerce ministry has reportedly begun consulting Alibaba, ByteDance, and Zhipu about restricting the transfer of their key training data out of China. It’s time policymakers in the U.S. treat our data ecosystem—which is much more prolific—with similar importance.
Regulate the Providers
Fortunately, the policy tools for regulating data vendors are already in place. The White House should simply build on the Department of Justice’s bulk data restrictions—a 2024 regulation barring the sale of Americans’ personal data to China—so that it covers AI training data. Direct sales to listed Chinese AI developers should be barred outright. For every other buyer, a know-your-customer burden would sit with the vendor: Before a sale, it would have to establish who actually owns the customer and what the data will be used to train—and it would be liable if it ignored obvious signs.
In February 2024, then-President Biden signed Executive Order 14117, which declared (under the International Emergency Economic Powers Act, or IEEPA) that “countries of concern” (including China) acquiring Americans’ bulk personal data is a national security threat. In fact, the order’s own findings warn that such data can “fuel the creation and refinement of artificial intelligence.” By April 2025, the Justice Department’s National Security Division had written the implementing regulation, called the Data Security Program, which prohibits or restricts classes of “covered data transactions.” Under the program, a data broker that sells Americans’ geolocation or genomic data in bulk to a Chinese entity faces outright prohibition, enforced by the Justice Department with civil penalties and criminal sanctions under IEEPA.
The problem, though, is that the Data Security Program applies only to personal data. Section 202.249 of the rule expressly excludes “data that does not relate to an individual, including such data that meets the definition of a ‘trade secret’,” thereby excluding much AI training data sold today from the bulk data restrictions.
Thus, the president should sign a new executive order declaring that the supply of AI development services to countries of concern is a national security emergency, and directing the Justice Department’s National Security Division to write a new rule for the Data Security Program.
The rule can draw on much of the same architecture for personal data: It could prohibit or restrict “covered AI development services,” with several enumerated data categories, from expert-crafted reasoning traces to reinforcement-learning environments. It could list “covered AI developers,” and delegate determination of data categories and covered developers to ongoing agency updates. Importantly, the rule should be aimed at the sale of data as a service, designed with customer specification; it would not apply to open-sourced environments or datasets available to the general public. This way, the rule can survive First Amendment concerns, and skirts the Berman Amendment to IEEPA, which carves out “informational materials” from direct regulation (a focus on the sale of data as a service is also the legal basis for the Department of Justice’s current bulk data restrictions).
The obvious objection to this proposal is that data is difficult to control. But this shouldn’t hinder policymakers. First, the data market makes enforcement tractable: Unlike an AI company that has to serve millions of people (making know-your-customer rules difficult to enforce), data providers serve only a handful of well-known customers. Second, we already limit the sale of data in other domains; there’s no reason AI should be the exception. Third, using IEEPA adds a private-enforcement layer: Existing whistleblower statutes (namely, 31 U.S.C. § 5323) pay 10 to 30 percent bounties from a $300 million pool for reporting IEEPA violations, which is particularly powerful in an industry dominated by a few firms employing tens of thousands of contractors each.
Lastly, while data flow to China will never be fully contained, a policy that raises Beijing’s costs is an improvement over the status quo, where the rule is that there is no rule at all. After all, chips get smuggled too. But the existence of smuggling isn’t a reason to start exporting our best chips.
The next Kimi moment is already in training. The most advanced U.S. chips shouldn’t power it. Let’s make sure U.S. data doesn’t either.
