Replacement Through Knowledge Acquisition
Certain species of wasps reproduce by laying eggs inside the bodies of caterpillars; when the eggs hatch, the larvae consume the caterpillars whole. In much the same way, frontier artificial intelligence (AI) labs are embedding their AI products deep within the operational organs of their most valuable enterprise customers.
Those customers should be concerned about what may one day hatch.
Just a handful of leading AI labs provide inference to millions of businesses across many sectors. These labs currently claim not to train their models on inputs from enterprise customers, yet they have strong incentives to alter or circumvent that commitment as their market power grows. By acquiring enterprise information, AI labs will be able to make their models more capable, widening the gap between private frontier AI and open-weight alternatives, furthering customer dependencies on frontier labs and, ultimately, developing the vertical expertise needed to compete directly with enterprise customers in their own markets—creating a deadly feedback loop. As it was with the land-grab training era of AI (and subsequent copyright litigation), by the time enterprise customers realize what is happening, it may be too late. Enterprises must protect themselves from “replacement through knowledge acquisition” (RKA)—from being consumed, digested, and replaced by frontier AI labs.
RKA Is a Threat to Enterprises
Replacement through knowledge acquisition occurs when an AI lab extracts information from an enterprise customer’s AI usage (for example, from prompts, uploaded files, uploaded data, model outputs, behavioral syntheses, and so on) and leverages that information to expand its market vertically, developing new products and services that compete with those of customers.
Given how murky the internal activities of AI labs are and how new the AI industry is, it is unclear whether or to what extent AI labs are currently engaging in RKA. However, the labs have a strong incentive to do so. New data, better models, and new markets will be critical to maintain the stratospheric growth that the market valuations of labs imply. Recent macroeconomic work on AI bubble dynamics helps formalize this pressure: High valuations can be sustained only with high growth and eventually high profits. Given that AI is a “general purpose technology,” there has already been a natural incentive for AI providers to integrate their models into a wide range of industries. The next step is for those providers to expand their market vertically to downstream uses, owning those uses, not merely servicing them.
The industries most vulnerable to RKA are those in which information shared with AI providers directly implicates the industries’ products or services—for instance, software or legal. In such industries, AI usage data from customers could most plausibly be employed to develop competing products or services. For example, employees of an enterprise software company might use AI models to write code or analyze data, disclosing critical information about enterprise software to AI labs. Similarly, in the legal industry, lawyers might employ AI models to execute legal workflows, revealing proprietary legal strategies. And in the financial services industry, AI models might reveal sensitive information about investment strategies. Moreover, if many employees within the same enterprise are using the same AI models simultaneously, the available information, once linked and synthesized, could reveal much more about the enterprise than any single employee has access to.
To make matters worse, there is no reason to think that the laptop-text-based interface for AI will remain predominant. AI startups are already building ambient note-taking devices, integrating agents directly into workflows, and incorporating visual AI into augmented reality products. Further in the future, embedded AI will escalate this effect—for example, deployed industrial robots can be used as data collection devices, extending the threat of RKA to industries like manufacturing, where incumbents may license robots only to have their industrial information appropriated by them. As the possible use-cases for AI expand, so too will the set of information available for acquisition and, thus, so too will the possible targets of vampiric inference. In other words, there may one day be no industry safe from the threat of RKA.
Pathways to RKA
There are a number of ways in which RKA could occur, differentiated along two key dimensions: (a) how conspicuous the lab is in its RKA activities—whether the customer’s information is acquired overtly or covertly, with the customer’s knowledge or without—and (b) the technical means a lab uses to achieve RKA—whether the learning happens through isolated training runs, requiring long-term data retention, or through continual learning, without a need for data retention. Below is a visualization of these two dimensions, outlining four possible modes of RKA.

On the first dimension, it is possible that enterprise customers may voluntarily consent to their information being retained or used by AI providers, especially where it doesn’t create direct legal risk or where AI providers offer compensation for information access. Alternatively, specialized models trained on enterprise information might become products, sold to the enterprises themselves. Rather than illicit data appropriation, enterprises thus might opt into data sharing agreements with AI providers to gain access to such specialized models.
However, where enterprise customers do not want to share their information, AI providers may still find ways to acquire it, either covertly—by violating or circumventing contractual terms—or overtly—by exerting market power to pressure enterprises into sharing.
Covertly, an inference provider might simply violate or circumvent its agreements. And although contract breach comes with legal risk, a calculating and sophisticated AI provider might decide that the risk is worth the reward, or they may attempt to creatively evade agreements by leveraging contractual gaps and gray areas. Note that covert acquisition does not necessarily entail illicit appropriation since, as discussed further below, it may be the case that an inadequate services agreement permits an AI lab to circumvent data usage restrictions. Overtly, an AI provider might obtain enterprise information from an unwilling customer by making use of growing market power to condition model access on more permissive information sharing, effectively coercing dependent enterprises.
In both cases, a lab would risk substantial reputational harm by engaging in RKA that is later discovered—indeed, suspicion that labs are engaging in RKA has already prompted rebukes across corporate America. But if a lab believes it can succeed at covert RKA, or if a lab has sufficient market power to engage in overt RKA, reputational risk may not be a sufficient deterrent.
On the second dimension, the two technical categories of model training are most importantly distinguished by their relative data retention requirement. Under the current paradigm of model development, models are upgraded stepwise, through gradual isolated training runs. New models are released as discrete versions (for example, GPT 5.1, 5.2, 5.3), and, once released, the weights for any particular version are relatively fixed. In other words, released models are not updated directly by user interactions. Through isolated training runs, engaging in RKA would require an AI provider to retain large amounts of user information to feed into subsequent training runs—a slow-loop form of learning.
However, RKA might become faster and more direct by using new methods of model training. One emerging example of this is continual learning: Rather than retaining large amounts of enterprise customer information to train future models, an AI provider might soon be able to leverage methods that allow models to be updated directly from transient interactions with customers, avoiding the need to retain information in a persistent dataset. In other words, as training methods evolve, the line between product use and model development will be blurred, and it will become easier for labs to leverage customer information without detection.
Across both dimensions, however, it is possible that other technical factors reduce the risk of RKA. For example, there may be structural limits to RKA: The knowledge that would allow an AI provider to compete in a new market might be tacit knowledge that can’t be pieced together retroactively from customer inputs. Or customers could adopt software architectures that make it more difficult to engage in RKA (for example, by using private clouds, open-weight models, or emerging frameworks such as unlinkable inference). And if several AI providers are competing at the frontier—or if open-weight models approach the performance of frontier models—AI labs may simply not have the market power to coerce enterprises.
Despite these contingencies, however, the unit economics of AI (that is, high up-front costs and low marginal costs to deployment) and the potential for technological flywheels such as recursive self-improvement mean that a future in which individual AI labs control vast market power—where open-weight and in-house models are not serviceable competitors to the private frontier—should be taken seriously. Already, as labs like OpenAI add new features to flagship products like ChatGPT, startups are being replaced by AI labs. And, more recently, the CEOs of larger enterprises—such as Microsoft’s Satya Nadella and Palantir’s Alex Karp—are waking up to the peril, warning companies of a growing concentration of power in a small number of AI labs. Reluctant enterprises will then be presented with a choice, Scylla or Charybdis: Embrace frontier models and risk RKA, or forgo the frontier and risk being outmatched by better-equipped competitors.
Legal Barriers to RKA
What legal measures can enterprises take to prevent or be compensated for RKA? There are three potential legal avenues to achieve this goal: contracts, trade secrets, and antitrust.
Contracts
Assuming that an enterprise wants to prevent RKA or be compensated for use of its data, services agreements that contain strong data security provisions are essential. But these agreements will contain strong provisions only if enterprises have enough bargaining power to demand them. If an AI provider becomes so dominant, and an enterprise so dependent, that the provider is able to require permissive data sharing provisions, then contracts will do little to prevent RKA. For example, the current leading frontier model from Anthropic, Fable 5, is already not available to organizations with zero data retention (ZDR) enabled. Although Anthropic claims that this retention is required for the sake of safety work, and that the retained data is not used to train models, the carve-out demonstrates that, under the right conditions, peerless models may enable AI providers to change information retention policies.
If the market power of leading AI labs continues to grow, their commitment to enterprise data security might erode. Notably, AI providers typically retain unilateral amendment rights and can easily reset terms at each new model release. Already, however, commercial AI agreements do not adequately protect unwilling enterprises from RKA. Though every major AI provider has promised to not train on enterprise customer data, their promises are vague and insufficient. For instance, the OpenAI enterprise services agreement claims only that “OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.” Anthropic’s commercial terms of service similarly claim that “Anthropic may not train models on Customer Content from Services.”
Both agreements, however, leave the scope of their covenants ambiguous. In OpenAI’s agreement, “Customer Content” is defined simply as “the Input and the Output,” which, in turn, are defined as what the “Customer and Customer’s End Users input to the Services” and “output from the Services based on the Input”—note the tautology of using the words “input” and “output” in their own definitions.
What it means to “develop or improve” OpenAI’s services is left unspecified. In Anthropic’s agreement, “Customer Content” is similarly defined, with “inputs” meaning “submissions to the Services by Customer or its Users” and “outputs” meaning “responses generated by the Services to Inputs.” What it means to “train a model” is also not specified.
These ambiguous protections create three issues for enterprises seeking to prevent RKA.
First, data breaches by AI providers are nearly impossible for enterprise customers to reliably identify because only the AI labs themselves know exactly what data is used for model training. Absent whistleblowing or transparency requirements, customers would need to somehow infer when a breach has occurred.
Second, since the scope of customer content is unclear, information shared by enterprise customers falling beyond the definition of “input” and “output” may still be used by the AI provider. Take Anthropic’s definition of output: “responses generated.” Does this mean that information must be a “response” to a user’s query to be an output? It seems possible that, under this definition, Anthropic could generate derivative, aggregated, or deidentified data from user inputs, not shown to the user as a “response,” which can then be retained for later use. Indeed, at least for the sake of monitoring model safety, Anthropic has even published methods for conversation distillation—that is, taking deidentified customer inputs and distilling them into categorized usage clusters. It is also unclear whether the scope of inputs—“submissions ... by Customers”—includes behavioral patterns, metadata, or other potentially valuable data that are not directly user “submissions.”
Third, vague contractual phrases such as “develop or improve Services” or narrow phrases such as “train models” could allow AI providers to engage in RKA activities that fall outside their scope. While these provisions likely still protect enterprises against both old and new methods of training AI models, other protections are less clear. For example, what if customer information is merely used to create evaluation benchmarks, tune system prompts, optimize routing, train separate classifiers (for example, for safety), or refine internal models? It is unclear whether such non-customer-facing activities would fall under OpenAI’s definition of “Services.” And what if an AI provider uses derivative information not constituting “customer content” (under current input/output definitions) to improve models?
In other words, certain AI development methods and data types might fall within contractual gray areas, a problem aggravated by how quickly advances in AI training techniques are being made. Enterprises should push to correct such ambiguities. The scope of “customer content” should be defined more broadly to include all enterprise information inputted, inferred, imputed, or synthesized from enterprise AI use, including any and all information distilled by AI models not shared as outputs with the customer. Similarly, rather than enumerating prohibited uses (for example, “Anthropic does not train models”), agreements should enumerate permitted uses (for instance, “Anthropic can monitor the number of tokens generated by models”), disallowing any non-enumerated activity.
However, even strengthened contractual provisions would not deter RKA if the AI labs expected the costs of breaching such provisions—that is, the enterprise’s remedial loss—to be lower than their expected benefits. Even a perfectly written contract fails at providing perfect deterrence. Thus, other legal avenues are necessary to consider.
Trade Secrets
Trade secrets might create barriers to RKA. But for RKA to constitute trade secret misappropriation, the enterprise information appropriated or used must be legally considered a trade secret, and the sharing of that information with a lab must not have destroyed that status.
Information qualifies as a trade secret if it meets two criteria: (a) the owner of the information has taken reasonable measures to keep it secret, and (b) the owner derives independent economic value—actual or potential—from the fact that the information is not generally known to or readily ascertainable through proper means by another person who can obtain economic value from its disclosure or use.
If enterprise information qualifies as a trade secret and is misappropriated by an AI provider (for example, by breaching a contractual duty to limit enterprise data use), or if misappropriation is merely threatened, then those enterprises may be able to sue for an injunction to prevent that misappropriation. A court might issue an affirmative injunction compelling acts (for instance, deleting the information) to protect the trade secret. Alternatively, a court might order labs to cease extractive practices or issue an injunction preventing the labs from engaging in threatened misappropriation or using misappropriated information.
However, there are several problems with relying on trade secret doctrine to prevent RKA. An initial hurdle is that an enterprise sharing their data with a lab must not have eliminated the trade secrets status of the underlying information. If an enterprise has shared trade secrets with an AI provider without express confidentiality requirements, it is possible its status as a secret will be destroyed. Assuming that is not at issue, the scope of any particular trade secret must still be specific—not “enterprise information” but rather “this piece of information.” A plaintiff cannot get trade secret protection by vaguely asserting that something within a corpus of their information is a trade secret. Their trade secret identification needs to be specific and particularized. This poses a problem in the case of vampiric inference, where information in question may be inherently vague and diffuse—millions of words of user prompts and files, billions of words of AI generated outputs, syntheses of user behavior, metadata, and so on. While certain pieces of information might still qualify, it is more likely that the true value of enterprise information is gestalt—an aggregated set of non-secret, generalized knowledge that is economically valuable not due to its secret status, but rather due to its specific application in context. While it is possible to protect compilations of publicly known information as trade secrets, the value of that compilation must still be due to its secrecy.
Even if a judge finds that an AI provider has appropriated a trade secret, if this secret is encoded directly into the weights of a model, it could be difficult to provide a remedy. While it is possible that a judge might order an entire model to be deleted, the disproportionality of such a remedy makes it unlikely (although, notably, alternatives such as forced royalties might still be available). While there has been some progress in the field of “machine unlearning”—the process of removing specific information from model weights—deleting information from a model is still far harder to do than deleting information from a database.
Trade secret doctrine might be a viable way to prevent certain pieces of information from being used by AI labs, but only in a narrow set of scenarios. And in scenarios where the leading AI labs have disproportionate market power, neither trade secrets nor contracts will be sufficient to deter overt RKA.
Antitrust
If AI providers concentrate market power, they may be able to condition access to their models on highly permissive data sharing and model training agreements. If those same providers then leverage that enterprise information to expand vertically into new markets, potentially competing with former customers, would antitrust provide a remedy?
There are two plausible moments at which an enterprise might turn to antitrust for protection from RKA. First, when the AI provider leverages its market power to condition access to models on information sharing (that is, exploitative abuse). And second, when the AI provider expands vertically, entering their customer’s market as a competitor (that is, monopoly leveraging). In both cases, however, it is unlikely that antitrust would provide adequate protections—after all, antitrust is designed to protect competition, not competitors.
At the first moment, conditioning customer access to an AI model on permissive data sharing agreements falls outside the scope of federal U.S. antitrust law. With regard to private business agreements, the Sherman Act polices only exclusion of rivals, not exploitation of customers, and thus contains no protections against RKA as a form of exploitative abuse. Unless an AI provider is using model access conditioning as a way to create or maintain a monopoly within the AI inference market itself—for example, by foreclosing rival AI providers using exclusive dealing arrangements that prevent enterprise customers from using competing AI products—such conduct is most likely to be understood by U.S. courts as well within the AI provider’s private right to choose its own business partners.
At the second moment, even if an AI provider later uses extracted enterprise customer information to enter that customer’s market, that entry would likely enhance competition in that downstream market, not reduce it. Only if the AI provider’s entry posed a dangerous probability of monopolizing the downstream market would there be a colorable antitrust claim. This claim might be strengthened if, after entering a downstream market, the AI provider continued to leverage its inference market power to extract information from its now-competitor. However, on its own, leveraging monopoly power in one market to compete in a second is not illegal. And a colorable antitrust claim would again require a dangerous probability of monopolization.
By contrast, in European law, exploitation alone is itself a cognizable offense. For example, during the EU’s investigation into Amazon Marketplace, Amazon was accused of using nonpublic business data from sellers to calibrate its own retail decisions and compete directly with the sellers—RKA in miniature. Amazon in response made a voluntary commitment to stop using this data, a commitment that was then codified into the Digital Markets Act (DMA). EU law, however, will do little to protect enterprises from RKA. First, the European Commission left AI services off the list of platform services regulated under the DMA. Second, EU antitrust proceedings can take years before authorities intervene, rendering any eventual remedies obsolete. And third, voluntary behavioral remedies are the norm in EU competition law and are often only partially implemented and poorly enforced.
Finally, state antitrust laws are also insufficient to remedy RKA. Although some state laws, like California’s Unfair Competition Law (UCL), are able to reach incipient antitrust violation, potentially helping enterprises enjoin coercive information extraction, these laws are remedially weak, and if information has already been acquired, model disgorgement would be an unlikely remedy. In general, state antitrust laws reflect federal law, and the same federal limitations to policing RKA will apply at the state level.
Therefore, under U.S. antitrust law, RKA creates a catch-22: Ex ante, the enterprise has no antitrust claim since the AI provider is not competing in its market; ex post, vertical expansion by the AI provider is merely competition on the merits. Ultimately, antitrust law is a dubious means of preventing RKA.
Regulatory Responses to RKA
It is difficult to know when or whether to regulate RKA. Made prematurely, RKA regulation could unnecessarily stifle AI innovation. Strict rules made to prevent inchoate RKA could be the equivalent of using wasp spray on a honeybee hive. It is possible that the benefits of RKA justify the costs—that is, creative destruction by competitive AI models may justify the downfall of incumbent enterprises.
Conversely, ignoring RKA entirely could be the equivalent of ignoring the wasp nest growing in your window A/C unit. Regulations that come too late could result in a large portion of the economy being controlled by a small number of AI labs, leaving only uncertain ex post legal remedies as a solution.
Therefore, our goal in this section is not to settle once and for all how RKA ought to be regulated, but rather to discuss potential regulations that could target it while maintaining policy optionality. In the short term, there are narrow measures that could be taken to discourage exploitative RKA. In the long term, if the threat of RKA escalates, more extreme measures might be required.
First, a problem that will pervade all efforts by enterprises to negotiate better contracts is an asymmetry of information. Model providers know which companies have negotiated strong contractual protections, but the companies that lack such protections don’t. This gives the labs an important informational advantage. Indeed, labs have historically been willing to offer customers better data protections when other customers’ protections become known. A straightforward reform would be to create a safe harbor for companies sharing negotiated data-protection terms among themselves. Such a provision might mirror the National Cooperative Research Act of 1984, passed in response to Japan’s increasing competitiveness in semiconductor manufacturing, which allowed U.S. semiconductor firms to collaborate on pre-competitive research and development without antitrust exposure.
Second, existing proposals for AI training data transparency and third-party auditing would enable better detection of RKA. There is a foundation for this already. Laws such as the EU AI Act require that providers of general-purpose models maintain technical documentation of their training processes for regulators. Transparency alone would help to resolve one of the fundamental issues with dealing with RKA: that only the labs themselves know exactly what information is being used to train their models.
Third, regulators could create a duty for AI providers to not misuse customer data. For example, AI providers could be prohibited from leveraging information derived from a customer’s AI inference usage absent express consent, codifying existing private commitments. Similar regulation has precedent in other sectors. In telecommunications, carriers may not use “customer proprietary network information”—that is, the knowledge they inevitably derive from carrying customers’ traffic—for their own marketing or other lines of business without consent.
And fourth, statutory damages could raise the costs of exploitative RKA (for example, by carving out punitive damages for data sharing breaches) where customers would otherwise be unable to secure sufficiently deterring damages. Similarly, states might add RKA-specific and per-violation statutory damages to the Uniform Trade Secrets Act, or expand the scope of trade secrets in the context of RKA. California would matter disproportionately in any such effort, and it has already enacted versions of training-data disclosure and limited per-consumer statutory damages.
* * *
Reliance on frontier AI providers is already a threat to the institutional sovereignty of enterprises. This reliance creates the preconditions for RKA. If a small number of AI labs continue to expand their lead in AI capabilities, the threat of RKA is only likely to increase. The existing legal tools are not well suited for addressing RKA. However, premature regulation carries the risk of stifling innovation. Therefore, the path now should be forward: warily watching the growing power of AI labs and laying the foundation for legal and regulatory responses to RKA if—or when—the time comes.
