The Most Important Phrase in GSA’s Revised AI Clause Has No Definition
The revision fixed much of the first draft’s overreach, but the clause still doesn’t say where the government’s data protections begin and end.
Over the past year, the federal government has drifted into governing artificial intelligence (AI) through its contracts. I have written extensively about why that model fails as a substitute for legislation: Bilateral agreements lack the accountability, deliberation, and permanence that statutes provide. Yet nodding vaguely to “procurement” has become the popular answer to AI governance—whatever Congress will not pass and agencies cannot mandate, the government can shape through its purchasing decisions.
Much of that discussion takes place at 30,000 feet, whereas contract clauses are written at ground level for a contracting officer, who must apply them to a real vendor, with real money at stake, and real government data moving through the vendor’s systems. An aspirational policy position can tolerate ambiguity—a contract clause cannot. Operationalizing AI governance requires drafting requirements clear enough to protect the government, definite enough for a vendor to price and perform, and workable enough for a contracting officer to apply tomorrow.
For decades, federal data-rights disputes have centered on files—what must be delivered; who may use, reproduce, or disclose them; and what happens to copies when the contract ends. That framework assumed the thing worth protecting was the stored file. AI inverts that assumption. Storage holds the files, while AI systems expose the underlying work: the questions asked, the drafts abandoned, the patterns of use, and the priorities they reveal. The General Services Administration (GSA), the federal government’s primary civilian purchasing agency, is now trying to address that inversion through a proposed clause that protects not only government content, but what government use reveals.
That is why, while much of the national AI debate centers on frontier safety, open-weight models, and strategic competition with China, I have spent the summer asking whether “government usage context”—a phrase that helps determine which information about the government’s work the clause protects—has a definition. It does not. Yet that phrase determines where the government’s data protections begin, where ordinary system operations end, and what vendors may do with what they learn from federal AI use.
A “Commercial” Clause
The GSA’s effort to adapt federal data protections to AI began earlier this year, when the agency issued the first draft of a proposed General Services Administration Acquisition Regulation (GSAR) clause for solicitations and contracts involving AI capabilities. In March, I described that draft in Lawfare as “governance by sledgehammer”—a clause reaching well beyond protecting the government’s data and into controlling the vendor’s product. GSA revised it substantially and, in June, published an updated draft, “Basic Safeguarding of Data within Large Language Model Artificial Intelligence Systems (LLMs),” that applies when LLMs process government data.
To the GSA’s credit, the June revision is a dramatic improvement and corrects much of the first draft’s most obvious overreach. GSA put down the sledgehammer and built a fence, without first locating the property line. That missing line predates the revision—several of the terms that determine whether government data protections apply were undefined in the first draft and remain undefined today. The result is a clause that restricts what vendors may retain and reuse without clarifying how far those restrictions extend.
That gap has significant consequences because most federal AI procurement runs through commercial acquisition: The government buys what the market already sells, largely on the terms the market already offers. Those terms are generous to the vendor. A commercial AI agreement typically permits the provider to retain operational information about how customers use the system, draw generalized lessons from that use, and fold what it learns back into the product it sells to everyone else. That is how these products improve, and it is much of what a vendor defends when it says a government requirement is “not commercial.”
The government does not have to accept those terms. GSA may decide that protecting government data and preventing a contractor from converting nonpublic federal insight into a commercial or competitive advantage justifies restricting what a vendor retains and reuses. But departing from commercial practice has a price. Government-unique restrictions may require separate systems, new controls, and commitments from their own suppliers.
Wherever GSA decides to draw those lines, it must draw them clearly. Contractors need to know what information they cannot use and what obligations they must price. Contracting officers need rules they can administer during performance. And agencies need to understand what protection they are buying and what functionality, competition, or value they may be giving up to obtain it.
The proposed clause raises more issues than a single Lawfare piece can cover. This article focuses on one recurring drafting defect in its data provisions: consequential restrictions whose triggers, limits, and proof of compliance are left undefined.
The Cost of an Undefined Boundary
The problems begin before the clause even applies. The clause covers solicitations and contracts involving the processing of “government data” by an LLM, with exceptions for LLMs embedded in common commercial products and for LLM functionality that is “incidental to the primary purpose of the core requirement being procured.” But “incidental” is undefined. A term that may determine whether the clause governs at all is left without a definition, factors, or examples to guide the contracting officer’s application.
Ordinarily, undefined terms acquire meaning through contract administration and disputes. That is a poor fit here for two reasons. The first is reach. This clause may be included in GSA contracting vehicles that move tens of billions of procurement dollars every year. And as the most visible AI clause proposed for government-wide vehicles, it is likely to become a template for other agencies, which is why the boundaries must be clear at the outset.
The second is application. An undefined boundary governing data use in an AI system does not sit dormant until a tribunal interprets it. It is applied continuously inside contractor and provider systems during performance. The contractor makes the first classification, and the government’s only window into it is a right to request existing documentation. But nothing requires the contractor to record that a classification was made, the rule it applied, or the records it excluded, and no one requests what they do not know happened. By the time the question is raised, the information may already have been retained, reused, or deleted.
Defining “Government Usage Context”
When someone uses an AI system, it generates two broad categories of information. One is the content of the interaction: the prompts, documents, other inputs the user supplies, and the outputs the system returns. The other is the operational and derived information generated around or from that use—information about how users interact with the system and what those interactions reveal about their work. The categories can overlap, but the first is what most people think of when someone says, “your data.” The second is where what I have called “informational advantage” often accumulates—the provider learns from how an organization uses the system, not only from the content it supplies.
That is why the familiar assurance “we don’t train on your data” promises far less than most people assume. It may prohibit only one use of specified inputs and outputs: training or fine-tuning a model. Standing alone, it says nothing about whether the provider may retain, analyze, aggregate, or use the “data dust”—the behavioral trail left behind from using an AI tool. A provider with access to those records can honor the “no training” promise and still learn a great deal from the pattern of use.
Now consider this in light of the sensitive information that can be derived from a government AI user. The GSA clause addresses this issue by defining “government data” to include “data outputs” such as logs, metadata, derivative data, and synthetic data. It gives the contractor only a narrow license to use government data for contract performance, support, and uses the contracting officer authorizes in writing, and separately prohibits contractors from using it for model improvement and specified business uses.
The “data outputs” definition creates an exclusion for technical system-level data, and that exclusion is essential. Without it, the definition could sweep in ordinary operational information a vendor needs to run the service. But the exclusion applies only if the data contains neither “government information” nor “government usage context,” and the clause fails to define either term. It offers only examples of data that may qualify: performance metrics, token counts, and processing times.
The exclusion has two defects. First, its own examples do not identify what is excluded. Those records fall outside the “data outputs” definition only if they carry no government information or usage context, so the same category of record can fall on either side of the line. Second, the clause is built as a broad rule with an escape hatch: The definition sweeps records in, the exclusion lets the contractor pull them back out, and the contractor makes that call first, inside its own systems. What is missing is the test—the point at which data about the system begins to reveal government information or use. You can run a truck through that gap.
Individually or in the aggregate, records such as these—task abandonment, refusal patterns, an activity spike at one agency, recurring use of a single feature—may reveal an agency’s workflow, capability gaps, priorities, or future demand. The classification is consequential. A record that satisfies the exclusion does not lose one protection. It falls entirely outside the data outputs definition, so many of the clause’s core data protections do not attach.
Of course, whether operational data reveals government use depends on the deployment, architecture, level of aggregation, and agency involved. But fact-sensitive does not mean standardless. Determining whether a price is “fair and reasonable” is also highly contextual, yet the Federal Acquisition Regulation (FAR) does not rely on the phrase alone; it supplies analytical frameworks for applying it. A clause intended for use across GSA’s government-wide vehicles likewise needs a functional test that can apply that discipline to more complex AI architectures.
Any workable definition of “government usage context” should answer three questions about the specific record or pattern, not the data category alone:
- Does this information exist in this form because the government used the system under this contract?
- Does it reveal how the government works (its activities, workloads, priorities, or methods) and not only how the system performed?
- If no single record reveals how the government works, do the records, in the aggregate, reasonably support a material inference about its nonpublic activities, workloads, priorities, or methods?
Applied to a planned data stream, the third question is a foreseeability screen. The contractor asks what the data stream is reasonably capable of revealing in this deployment at this level of aggregation. This is similar to how GSA runs its own flagship AI deployment, USAi, which distinguishes among types of system data and applies different retention and use rules established in advance.
The test also exposes two policy choices GSA must still make. Does “government usage context” refer to use by the ordering agency or by the government as a whole? Records that reveal little at the government-wide level may be revealing when tied to a single agency. And how demanding should the proposed materiality threshold be—does any inference about government use qualify, or only one that is operationally significant or capable of conferring a competitive advantage?
The drafting solution is straightforward: define “government usage context” through the test. The examples GSA currently provides should illustrate the test, but they cannot substitute for it. And because the contractor classifies records first, GSA should also specify what showing is required if the government challenges a classification.
Downstream Use of “Government Data”
Defining “government usage context” helps determine which operational records fall under the government data regime. It does not settle what the vendor may do with them afterward. That is a harder question, because the answer changes as the connection to the original data becomes more attenuated.
The clause makes a comprehensive choice: Government data may be used only for performance, support, and any other use authorized in writing by the contracting officer. Everything else, including model improvement, is prohibited. That is a protective default, but it does not specify when applying a generalized lesson remains a use of government data.
The easy cases are, well, easy. Fixing an outage as part of a support requirement under the contract is permitted. Using government data about agency demand to shape a sales strategy is prohibited. But consider a harder case. The contract requires the vendor to study where government users abandon a workflow and improve that workflow for the agency. The abandonment data are government data, and a written finding derived from them may be as well. Using either to improve the vendor’s commercial model would be prohibited.
However, during that permitted work, the vendor may also learn a broader design principle that contains no government-specific information and may be supported by its experience with other customers. If the vendor later applies that principle elsewhere, is it still using government data? Or has the connection become too attenuated? The clause does not say.
The contracting officer-authorization mechanism does not solve the problem. It governs what happens after a proposed use has been classified as outside the permitted purposes. It does not tell the parties when attenuation has occurred. Until the clause supplies that rule, neither the vendor’s compliance program nor the contracting officer has a standard for determining which later uses require authorization.
The GSA does not need to police what engineers remember. It needs an administrable rule, and procurement law already handles a version of this problem. FAR subpart 9.5 requires contracting officers to identify and evaluate potential organizational conflicts of interest (OCIs)—conflicts arising from a contractor’s other work or relationships that may impair its objectivity or create an unfair competitive advantage—and to avoid, neutralize, or mitigate potential conflicts before award. One type of conflict is an “unequal access to information” OCI, in which a firm, through contract performance, gains access to nonpublic, competitively useful information that may give it an unfair advantage in a future federal competition. The inquiry is fact-sensitive but administrable: Are there hard facts that the firm had access to nonpublic, competitively useful information, and did any mitigation adequately address the potential advantage?
Drawing from this OCI framework, an overlap between personnel or systems with access to covered government data and those engaged in general product improvement or other downstream commercial work could support an inference of prohibited use. A contractor could then rebut that inference by showing independent development through contemporaneous documentation, dated product road maps, firewalled teams, or a clean-room process.
There are limits, however, to the OCI framework, because informational advantage is broader than an unequal-access OCI, which requires a nexus to a later federal competition. Informational advantage can accumulate even when no later procurement is on the horizon, and its value may lie in commercial product development that OCI doctrine does not reach standing alone. The clause should borrow the doctrine’s framework (access, segregation, and mitigation) and supply what the doctrine does not: a rule for when a generalized lesson remains a use of government data, and a showing of independent development sufficient to rebut an inference of prohibited use.
GSA is close to that structure. The clause already requires access controls and segregation of government data from other customers’ data. But it does not require separation between government support work and broader product development or explain what role such controls play in establishing compliance with the downstream-use prohibition. The drafting distance may be small, but the impact on governance is not.
The clause should therefore define four things: (a) when applying a generalized lesson remains a prohibited use of government data; (b) what access controls or mitigation establish compliance; (c) when a later use requires written authorization; and (d) what showing is sufficient to establish independent development. GSA may choose strict containment or define a controlled-reuse path, but it cannot leave contractors and contracting officers to infer which regime the clause requires.
From Aspiration to Obligation
The June revision shows that GSA can distinguish between protecting the government and controlling the vendor’s product. The next revision must do something more difficult: distinguish an enforceable protection from an aspiration.
This is the part of AI governance that no one writes about. Procurement is treated as the back office of AI policy—the place where big ideas are sent to be implemented. This clause is what implementation looks like. As I said in my last piece, I do not envy the people writing this clause. Procurement is a constant series of trade-offs—precision against flexibility, the government’s interest against a market it needs to keep in the room—and both are present here at once, in a technology none of us fully understands. GSA gets to decide how to strike those balances. But you cannot strike a balance between interests you have not defined.
