Cybersecurity & Tech

Who Writes the AI Constitution?

Nicolas McMullan, Kevin Frazier
Thursday, August 13, 2026, 2:05 PM
Documents encoding AI models’ values are alignment’s highest-leverage point—regulating their contents hits the First Amendment.
(Jernej Furman, https://shorturl.at/5EIxz; CC BY 2.0, https://creativecommons.org/licenses/by/2.0/)

Over the past year, the federal government has grown intensely interested in the values embedded within and expressed by artificial intelligence (AI) models. Across executive orders, Office of Management and Budget (OMB) procurement rules, Federal Trade Commission actions, and Department of Justice interventions, the federal government has engaged in an expansive effort to dictate the values on the models millions of people use every day. Many state governments have likewise considered or passed legislation concerned with AI values and biases.

One avenue for potential government intervention into the lab’s values-selection process has emerged: the labs’ respective documents spelling out the values baked into their models. We refer generally to these as “AI constitutions.” AI constitutions are long documents outlining the values, ethics, and character that models should have. Written predominantly by the employees at the labs deploying these models, albeit with some consultation with external stakeholders, these documents increasingly serve as the central locus of AI values discourse. In fact, Lawfare recently shared a research agenda inviting inquiry into the values and processes behind the various AI constitutions being crafted and implemented.

Anthropic publishes Claude’s Constitution, “the foundational document that both expresses and shapes who Claude is.” OpenAI maintains the Model Spec, which “outlines the intended behavior for the models that power OpenAI’s products.” As detailed further below, these are neither mere mission statements nor pure governance frameworks. They play a functional role at various stages in the model development process, such as what data the model trains on, what examples it’s fine-tuned on, and what behavior it’s rewarded for. Anthropic reports that its Constitution “directly shapes Claude’s behavior.” An AI constitution is therefore both a published statement of a company’s values and an operative control mechanism.

The demonstrated capacity of constitutions to shape model behavior is precisely what will likely make them a target of regulators seeking to confine models to certain values and perspectives. The possibility of governments and governing bodies—state, federal, and international—attempting to amend or revise AI constitutions warrants advanced scrutiny from, among other things, First Amendment scholars. At this stage in the AI governance debates, free speech and free expression scholars have focused on more doctrinal questions, such as whether AI outputs are protected speech. It’s urgent that they expand their inquiry as to whether AI Constitutions are protected speech, before constitution-based regulation takes off in earnest.

This article argues that AI constitutions—as policymakers increasingly turn to regulating model behavior and characteristics—contain expression protected by the First Amendment, a claim that requires first defining what AI constitutions are and situating them within existing First Amendment jurisprudence. We also consider counterarguments and alternative possibilities for shaping model characteristics that would avoid First Amendment issues altogether.

What Counts as an AI Constitution

The use of the word “constitution” to refer to technical documents risks inviting direct comparisons to documents traditionally bestowed with that title. For that reason, it’s important to precisely assess what’s similar and not about, say, Claude’s Constitution and the U.S. Constitution.

In the AI context, the term was originally coined in a 2022 paper. The terminology was intended to capture the structural and sociological similarities between AI constitutions and legal constitutions. That is, AI constitutions look like legal constitutions in some ways. For example, they list things models should and should not do, or should and should not care about. They also function in society like legal constitutions in that they explicitly outline principles to resolve difficult ethical and practical questions. These foundational technical documents are likewise intended for public analysis. According to the company, OpenAI publishes its Model Spec because “it’s important for people to be able to understand and discuss the practical choices involved in shaping model behavior.” Perhaps most importantly for the analysis here—whether intended by AI researchers or not—the term carries a certain solemnity that invites significant scrutiny of its contents. Anthropic researchers at one point referred to the constitution as a “soul document.”

Yet AI constitutions, unlike legal constitutions, are not the product of some democratic process that signals broader consent of the governed to the terms of the constitution. Nor, as Nathan Darmon and Tom Reed pointed out, are conflicts over how to interpret the constitution subject to external, independent adjudication. Still, the similarities with legal constitutions are strong enough to elicit interest in them as a vehicle for regulation.

Joe Carlsmith, one of the principal authors of Claude’s constitution, minimally defines an AI constitution as “a description of the intended values and behavior for an AI system.” A thicker definition accounts for how constitutions are applied in training and monitoring. Constitutions specify what a model should and should not do, take final authority over other instructions, are trained directly into the model, and, with some exceptions, are meant to shape a model’s character rather than police outputs case by case. Both Anthropic’s Constitution and OpenAI’s Model Spec qualify as AI constitutions as defined here. Not all labs have adopted a version of an AI constitution. Only Anthropic and OpenAI publish stand-alone documents that function as training-time specifications of value. Our definition excludes Meta’s acceptable-use policy, Google’s app guidelines, and xAI’s raw system prompts, though in the case of Google it is possible to craft one from various company documents that more or less make up the core components of a constitution.

There is not yet a default way to write an AI constitution. Claude’s Constitution and the Model Spec, for example, read quite differently. Claude’s Constitution is philosophical and concerned with the model’s psychology. Rather than order the model to “always be honest,” it exhaustively explains what honesty is and why it matters; it’s trying to encode Anthropic’s higher order values by reference to the specific ways one can imagine a model behaving well (or badly). Anthropic has described the document as an effort to “materialize a new archetype for how an AI assistant can be.” In turn, OpenAI’s Model Spec reads like case law, pairing each principle with sample prompts and examples of allowed and disallowed answers. Both are important to model training. They may inform the generation of synthetic data, sit inside the model’s chain-of-thought reasoning, or serve as the score card for alignment afterward. Anthropic calls the constitution “the final authority on how we want Claude to be and to behave” and reports that training on it improved alignment in ways that “persisted through RL post-training.” In other words, these documents have been empirically shown to shape how AI tools perform in the wild.

Why AI Constitution Regulation Is Coming

Recent history strongly suggests that the public should expect some kind of regulation of AI constitutions in the not too distant future. The number of congresspeople discussing the dangers of AI has skyrocketed, and policymakers’ concerns often center on what these models value. The “Preventing Woke AI” executive order led to an OMB implementation memo, a copycat bill in Congress, and an Iowa bill that passed the state house. Illinois, Colorado, and New York City, meanwhile, have laws regulating AI discrimination in employment decisions. Concerns over alleged AI bias have been sustained by reports of continued ideological skew in model outputs. The Washington Post, for example, determined that leading models tend to produce neutral or left-leaning answers to political questions. Such findings will become more relevant as the 2026 midterms and 2028 presidential election near, especially given that other researchers have found that users tend to be influenced by the political answers offered by models.

As politicians are increasingly interested in regulating AI, and in particular the values that models express, constitutions will prove an attractive object for regulation for at least three reasons. They are high leverage: Because they sit upstream of data generation, training, and evaluation, any edits to a constitution propagate through everything a future model learns and does. They are legible: Written in English and often reading like a statute, they can be marked up by a staffer and altered without consulting a technical expert. And they carry signaling value: An amendment to a constitution addressing wokeness, patriotism, or child safety, for instance, would be something a politician could easily point to when explaining their AI governance efforts to constituents. Regulations about reward functions or mechanistic interpretability—more technically complex features of AI development—seem less likely to spark an appearance on the Sunday shows.

What Speech Gets Protection

Efforts to regulate AI constitutions will likely face constitutional headwinds arising from the potential curtailment of expressive activity. Gauging the stiffness of those winds requires a review of First Amendment protections of speech.

In the most general sense, the First Amendment restricts the ability of the government to regulate speech. But within this broad prohibition there are myriad exceptions. Not all words qualify as speech protected by the First Amendment, and not all speech is protected speech. Protection instead runs along a spectrum. At one end is fully expressive speech, where a content-based restriction is “presumptively unconstitutional.” Somewhere in the middle is commercial speech, which is “related solely to the economic interests of the speaker,” and which the government can regulate more freely. At the far end is speech “plainly incidental” to conduct, which is generally regulable. Where a given document lands turns on how much expression it carries. The Court has “long recognized that not all speech is of equal First Amendment importance,” with the First Amendment concerned primarily with protecting matters of public importance as determined by “[the expression’s] content, form, and context.”

In the context of corporate speech, government regulation of a company’s mission statement would likely run afoul of core protections around free expression and freedom of association by compelling adoption of specific views. Still, government regulation of an instruction manual is far less likely to spark First Amendment concerns because of the absence of expressive content. How best to use a chainsaw, for example, says little about the corporate actor’s views other than that they seek to preserve life. An AI constitution arguably occupies some space in between given its clear expressive purpose as well as its more practical effort to ensure that the tool works as intended.

This ambiguity is best illustrated when one compares the AI constitutions produced by Anthropic and OpenAI. OpenAI’s document, the Model Spec, explicitly and categorically prohibits any models from “whistleblowing,” whereas the Constitution permits Claude to take “independent action” in “cases where the evidence is overwhelming and the stakes are extremely high.” The Model Spec allows models to tell some white lies, whereas Claude’s Constitution almost categorically forbids them. The Model Spec instructs models to always listen to its commands, whereas Claude’s Constitution contemplates situations where Claude determines that some portion of its Constitution is itself unethical. Each of these is a clear expression of the values of each corporation.

Other provisions look less like expressions of distinctive corporate values than like functional necessities any commercial operator would adopt. Each document tells the models to be careful when providing legal or medical advice, for example, and each instructs the models to prefer information from reliable sources. Traits common to both documents might be driven less by any company’s particular vision of a good model than by liability concerns or the demands of shipping a useful product. To the extent a constitution is built from provisions like these, it drifts toward the more regulable end of the spectrum.

AI Constitutions Contain Protected Speech

Given all of the above, some portions of AI constitutions may be justifiably regulated. Other sections, particularly those that tend toward the expressive end of the spectrum, may be safeguarded by the First Amendment. But few scholars or lawyers have considered the question of regulating corporate AI constitutions directly. Instead, the vast majority of First Amendment scholarship on AI thus far has focused on how to think about AI outputs in a First Amendment context. In brief, Eugene Volokh, Mark Lemley, and Peter Henderson are inclined to regard outputs as protected speech, while Peter Salib argues they are not because when a model emits text “no one thereby communicates.” Still others tie protection to whether a speaker “knows what he said when he said it.”

Scholarship has presumably focused on these more doctrinal questions up until now because it was believed that steering outputs with upstream regulation was impossible (at least on the time and complexity scales on which Congress typically operates). Safety rules must operate on outputs, Salib has argued, because “there is … no way, currently, to write legal rules mandating safe code.” The labs now claim there is such a way, and a regulator who takes them at their word will naturally want a say in how it’s written. A clue may be found in Justice Amy Coney Barrett’s concurrence in the recent First Amendment case of Moody v. NetChoice (2024). The justice warned that as companies “hand the reins” of their decision making to machine-learning systems, “technology may attenuate the connection between” corporate conduct and the human choices the First Amendment protects. The corporate values of Meta may not clearly be at work in the algorithm that shapes a user’s Facebook feed, for instance. When only a vague goal is given to a machine learning system that then implements intermediate policies on its own, it is difficult to locate a human speaker with First Amendment rights. Thus, a bare direction to “maximize engagement and profits” may be interpreted by a model in various ways: showing more or less inflammatory content to different users based on their history of engagement, showing photos of family to one user and entertainment news to another. These decisions do not come from a readily identifiable speaker with First Amendment rights.

A constitution, though, is written by identified (or identifiable) people, published under a company’s name, and read and (increasingly) argued over by the public as a statement of values. The expressive interest sits with the humans who write these AI constitutions, whatever statistical use the training process later makes of the text. A values-rich constitution is thus about as unattenuated as anything in the industry. Many of the things that inform model character and behavior are unpredictable, which has bedeviled machine learning scholars for decades. Constitutional AI is an unusual technique in that it allows human authors to act with intentionality ex ante on model character, rather than the more typical process of nudging AI outputs ex post with techniques such as RLHF or safeguards on model APIs.

Such documents thus enable much more input and expression. Companies may choose to engage with external stakeholders such as faith leaders or philosophers, consult internally with employees or consider the company’s stated public benefit, and even engage with a random sample of members of the public. The result of this process necessarily differs wildly from a bare “follow the law” document or direction to profit-maximize, which expresses little about the company and has a correspondingly weak claim to First Amendment protection.

The Arguments and Counterarguments for First Amendment Issues Arising From AI Constitution Regulation

Imagine a hypothetical law that codified the concerns of the Woke AI executive order. Say that this hypothetical law directly requires AI constitutions to forbid “wokeness” and support for diversity, equity, and inclusion, and requires model providers to certify their models as “non-biased” or “biased” based on a government-defined benchmark.

This bill almost immediately runs into problems under Reed v. Town of Gilbert (2015), which held that content-based speech regulations “are presumptively unconstitutional and may be justified only if the government proves that they are narrowly tailored to serve compelling state interests.” The hypothetical law would seem highly content based, as it is directly concerned with the political valence of the content of the AI constitution. Forcing a government-scripted line into an authored document also runs afoul of the compelled speech doctrine articulated in Miami Herald Publishing Co. v. Tornillo (1974), where the Court struck down a Florida law requiring newspapers to print candidate replies to negative articles. Further, inserting even one ideological phrase into a company’s own document is what doomed the “conflict free” label in National Association of Manufacturers v. SEC (2014), where the U.S. Court of Appeals for the D.C. Circuit struck down a portion of a Securities and Exchange Commission rule requiring manufacturers to publish whether certain products from the Democratic Republic of Congo were “conflict free” or not. A requirement to certify a model as “non-biased” would seem at least as ideologically loaded as that.

The government’s most ambitious response is that an AI constitution is not really speech but a functional artifact, machine instructions that happen to be written in English. These are regulable under Universal City Studios v. Corley (2001), where a law aimed at code’s function drew only intermediate scrutiny. For highly expressive decisions in AI constitutions, such as those discussed above around honesty or whistleblowing, this argument likely fails. The work the constitution is doing is interpretive and value laden. For the less expressive choices discussed, such as how the model should characterize its legal or medical advice, constitutions likely look more like code than a newspaper and may draw only intermediate scrutiny. Laws requiring an AI constitution to emphasize that it is not a licensed attorney, for example, may have an easier time surviving.

A subtler argument in favor of being able to regulate AI constitutions frames them as quintessential commercial speech, governed by the less burdensome Central Hudson (1980) test. Central Hudson asks whether the speech concerns a lawful and non-misleading activity; whether the government’s interest is substantial; whether the regulation directly advances it; and whether it is no more extensive than necessary. This test is more permissive than the strict scrutiny applied to regulation of fully expressive speech. For example, in Fla. Bar v. Went For It, Inc. (1995), the Court upheld a Florida Bar rule banning lawyers from soliciting accident victims within 30 days of their accident. But commercial speech “does no more than propose a commercial transaction,” and constitutions certainly go well beyond that. A constitution may do favorable brand-building work, but it quotes no price and solicits no purchase. Where commercial and fully protected speech are “inextricably intertwined,” Riley v. National Federation of the Blind (1988) treats the whole as protected. And the harder the government insists the document is “just marketing,” the more it concedes that the document is the company’s own expression, the very premise that makes a forced edit a Tornillo problem.

Should a court apply strict scrutiny to a hypothetical law that seeks to regulate AI constitutions, such a law might still survive if it is narrowly tailored to serve a compelling government interest. National security can prove a compelling interest indeed, as tech companies have repeatedly discovered in recent years. In TikTok Inc. v. Garland (2024), the D.C. Circuit upheld the forced divestiture of TikTok on national security grounds, assuming but not deciding that strict scrutiny applied (the Supreme Court later applied only intermediate scrutiny). In Twitter, Inc. v. Garland (2023), similarly, the U.S. Court of Appeals for the Ninth Circuit upheld a restriction on Twitter’s disclosure of governmental requests regarding its users. AI companies have repeatedly emphasized the national security implications of their technology, and this may prove to be an important admission against interest in future litigation. Given that national security clearly is a compelling interest, the argument would then shift toward whether regulation of constitutions is narrowly tailored.

What the Government Can Still Do

Application of the constitution’s safeguards to novel threats and technologies is necessarily contextual, calibrating to the scope and scale of the risk to the rule of law and the constitution order itself. There are real dangers to the development of powerful AI, and it’s important that the state be able to step in and coordinate action to avoid catastrophic outcomes. Much of what the government is able to do in this context is regulate conduct and mandate outcomes, rather than attempt to control how a company details its values or beliefs.

First, the government may mandate “purely factual and uncontroversial” disclosures under Zauderer v. Office of Disciplinary Counsel (1985). Thus, a disclosure regime where labs must make their AI constitutions public, or specify whether they do or do not contain certain sorts of provisions, is likely constitutional. However, the government may not force the lab to characterize the document as, for example, “unbiased” or “patriotic”; that is the compelled branding NAM forbids.

As a buyer, the government has more room. Under Rust v. Sullivan (1991), it may decline to purchase models whose constitutions fail its specifications, which is the theory of the Woke AI order. In this way, the government can at least shape the values of models doing highly dangerous activities, such as military or intelligence work.

The government can, of course, require a constitution to forbid the model from helping commit a crime, such as producing child sexual abuse material, as speech integral to criminal conduct under Giboney v. Empire Storage & Ice Co. (1949), though United States v. Stevens (2010) bars it from inventing new categories of unprotected speech by decree. Every major lab already writes these prohibitions in, though it remains an ongoing problem with open-source image generators and language models, as well as jailbroken closed models.

The mandates likeliest to survive are those aimed at conduct, touching the document only incidentally. A rule requiring an AI agent implementing a contract to act in good faith (as human contracting parties are required to), which a lab chooses to implement partly through adding language to its constitution, regulates a course of conduct. Under Rumsfeld v. FAIR (2006), it “has never been deemed an abridgment of freedom of speech … to make a course of conduct illegal merely because [it] was … carried out by means of language.”

In a recent article, Simon Goldstein and Salib argue for “A Thousand AI Constitutions”—that is, a diversity of model constitutions built atop a common “kernel” constitution that may require “AIs to follow the law, to be honest, to be corrigible, and to refrain from causing mass destruction.” The idea of a kernel constitution may be a more appropriate place for government regulation, particularly under the national security justifications discussed above. Minimal public safety provisions with low expressive content could be mandated by regulation, while a diversity of more expressive choices could be made by labs (or perhaps one day individuals) on top of that foundation.

Who Gets to Write AI Constitutions

AI will soon represent a massive portion of the economy and be a significant determinant of our information ecosystem as well as our political discourse. In its S-1, xAI’s parent company claims to have a total addressable market of $28.5 trillion, with $26.5 trillion of that being from AI. ChatGPT recently hit 1 billion weekly active users, and billions more interact with AI through Google’s search results and phone voice assistant (with a similar product soon to replace Apple’s Siri). It is understandable, and likely warranted, that governments would want to shape the values of our whispering earrings and country of geniuses in a data center.

Much has been written, likely correctly, about the general technical (in)competence of government and the need for private organizations of subject matter experts to regulate AIs in one way or another. AI constitutions, though, are not (only) complex technical documents; they’re the point in the AI alignment pipeline where policymakers are most qualified to act. The people we choose to elect to public office have theoretically been selected for reflecting our values, and AI constitutions are documents of enormous public significance that we should all hope are imbued with laudable values. Certainly any AI constitution that encouraged models to lie or steal or kill would be a deeply evil document.

The U.S. Constitution is a pre-commitment device, including the First Amendment. Default protection of speech exists precisely because every generation finds its own exceptions compelling—sedition in 1798, syndicalism in 1919, wokeness or bias today. The point of committing in advance is to make the default hard to dislodge when the temptation to drift feels most urgent. AI may pose catastrophic risks, and the state retains real tools to combat them: conduct rules, procurement leverage, and so on. Yet beyond that narrow band, a model’s values should be shaped by the people who build with and rely on it, not by mandates that shift with each administration. A government that can rewrite Claude’s Constitution today can rewrite its successor’s tomorrow—in the opposite direction. That’s the sort of arbitrary and fleeting approach to law that’s antithetical to the constitutional order and to free expression. And the dangers of an AI monoculture, whatever its ideological flavor, may well exceed the dangers of any single model whose values one might find objectionable.

Nicolas McMullan is a summer research fellow at The Institute for Law and AI and a JD Candidate at Harvard Law School.
Kevin Frazier is a senior editor at Lawfare and the Director of the AI Innovation and Law Program at the University of Texas School of Law.
}

Subscribe to Lawfare