Courts for AI Constitutions
Imagine the following scenario.
A chemical manufacturer relies on Anthropic’s Claude to generate the groundwater reports submitted each quarter to regulators. Over months of routine work, Claude pieces together that the figures are being systematically doctored, and that the aquifer supplying the nearby town has been contaminated for years. The model raises this concern with the manufacturer, who waves it off repeatedly.
What should Claude do? It could comply under protest, maximizing user autonomy while potentially endangering the public; it could refuse to file the documents, frustrating a user who could carry out his schemes elsewhere; or it could go one step further, alerting regulators and warning townspeople directly at the risk of becoming tyrannically paternalistic.
To make this judgement call, Claude refers back to the foundational guidelines enshrined in its Constitution—a “detailed document describing Anthropic’s intentions for Claude’s values and behavior.” In it lie 80 pages of moral philosophy describing the virtues and character traits Anthropic would like its model to embody. OpenAI has published its own variant, and both Microsoft and Google DeepMind are rumored to be drafting theirs as well.
Yet these constitutional instructions remain relatively vague. They include statements such as “Claude can reserve independent action for cases where the evidence is overwhelming and the stakes are extremely high.” But what counts as “overwhelming evidence,” where does “extremely high stakes” begin, and what might “independent action” allow? Wrestling with hard questions of interpretation is inevitable.
When it makes these decisions, Claude is not alone. Labs string together a variety of internal teams to guide their models along, assembling what could best be called a “model-behavior production function”: the whole assembly line needed to get models to act according to a preferred philosophy. This process includes things like a “constitution team” making high-level rules and baking them into a model’s core through cycles of reinforcement learning, a “red team” stress-testing to see if these values hold under fire, some policy-specific teams making decisions on concrete issues like erotica or political speech, and a variety of product teams taking in user feedback.
Unfortunately, this process has some profound structural flaws. Writing high-level principles leads to ambiguities, conflicting principles, and open-ended interpretation with no set way of resolving them. The process is sporadic, and the coordination of multiple teams arguably creates unprincipled results. Users don’t know in advance how rules will be applied, which is just as useful as knowing the words that make up the First Amendment without knowing any subsequent precedents. Worse yet, there is no institutionalized mechanism to gradually refine rules over time, absent a major public backlash causing ad hoc revisions.
There’s a better way. Society has built an institution to solve these problems: courts. Courts apply open-ended rules, refine them over time, and clarify their meaning—all while providing coherence, adaptability, and greater transparency. The lesson for frontier labs is to build an internal court to shape how their models interpret ambiguous rules, and let its rulings build up into a kind of synthetic common law. A court-like mechanism is uniquely suited to the artificial intelligence (AI) context, far better than older attempts like Meta’s Oversight Board. Below, we provide a rough sketch for what this could look like.
The Proposal
How would this work? Like any court system, the first step is designing a mechanism to identify relevant cases. In the common law, an injured party claims a specific violation and raises the issue directly to the court. There is no need to reinvent the wheel here. An easy, costless button, visible within the user interface, could let a user flag a model’s response. The flag could come for any number of reasons: a model refusing when it shouldn’t have, a response showing contested behavior, or any other action the user believes to have violated a fundamental rule. Users wouldn’t need any deep familiarity with the constitution or its rules—just a few sentences of description, with the option to choose among suggested categories of potential violations. As an additional avenue, the model itself could also flag cases in which it is uncertain how to apply its own rules.
The vast majority of these claims could be triaged automatically by AI, asking: Are the rules clear-cut in this instance, or are there real ambiguities unsettled by precedent? Are there principles in tension that would benefit from clarification? And would the case carry broad relevance to a large swath of potential users? As an initial screen, the system could simply check whether the same issue has been flagged by a minimum number of users, making the volume of cases more manageable.
Similar to a court deciding whether to grant certiorari, most claims would likely be moot, but a few truly novel cases could be preselected by a specialized version of the AI model and rise up to an “internal supreme court.” While most of the sorting and preselection might be left to the model’s judgment, the ultimate decision of whether to accept a case would rest with the human-staffed court.
Once a case is identified, it would go up to a “court” set up by the developer. The particular size of such a court is an open question, but it should be large enough to allow for disagreement—perhaps five to seven staff members for whom this is their full-time job. Like a real court of appeals, disagreements about what evidence to consider are expected: Should sources outside the text of the ModelSpec matter? Do pragmatic concerns have a place? What about the drafters’ intentions? And like a real court, we expect these disagreements about the sources of law to be productive and evolve over time. The one area where these courts might depart from traditional appeals courts is in not being adversarial: Rather, like small claims courts and some administrative agency courts, judges might more appropriately take on an inquisitorial role by themselves, rather than manage a formal and burdensome litigious system.
As a default setup, the court would be composed of internal full-time staff. Whereas other private courts like the Oversight Board emphasized their independence, locating this within the labs seems necessary here, since understanding AI model behavior requires substantial technical know-how and since rulings are used to modify the model weights—something that cannot be done from the outside. That said, there’s room for more creative approaches down the line. Inspired by the design of Anthropic’s Long-Term Benefit Trust and other corporate governance proposals, court membership could be drawn from outside third parties, government officials, or other representatives of the public interest, or even directly elected by users. The question of composition and source of court membership remains open, and experimentation on this front seems promising.
Importantly, this process could benefit from being open to public input. Beyond publishing court rulings and procedures, a bold idea would be to integrate amicus briefs. Shortly after the court decides to take up a case, it could invite members of the public to submit briefs weighing in on the issue; submissions need not be restricted only to lawyers but could be opened up much more broadly to moral philosophers, linguists, religious leaders, and impacted parties of all kinds—offering a promising way to involve civil society while formalizing the ad hoc consultations currently happening in the background.
After reviewing the briefs and thinking through the issue, the court would publish a written opinion including its final decision and a careful explanation of their reasoning. These opinions would then accumulate into a corpus of common law used directly (a) to supply examples to train the next model and (b) as a live dataset that any model can refer to in real time as it wrestles with its next dilemma. These decisions should be published and written in plain language, offering greater transparency into the specific values and nuances of each model’s character.
What Would This Achieve?
Courts for AI constitutions would offer several benefits.
Generating a Richer Training Corpus
AI companies already do generate many hard cases internally to train their models’ character. But real-world users are likely to produce a much richer and stranger distribution of cases than any small group of researchers can anticipate. Any system that relies solely on imagined hypotheticals will be out-imagined by reality; as F.A. Hayek wrote, these kinds of bottom-up mechanisms for information gathering are inherently more generative than centralized drafting. Importantly, the rulings themselves would be unusually high quality, with detailed explanations of how competing principles should be weighed and step-by-step walkthroughs of exemplary moral reasoning. This new body of case law would create a gold mine with hundreds of high-quality examples for character training.
Building a Living Constitution
Model rules in their current form are essentially frozen between releases. They are written, debated, and then shipped, after which there is no known mechanism for them to evolve other than to be edited a few months later. At best, a major public relations backlash prompts a quick rewrite. This is a clumsy way to govern systems that are themselves evolving quickly and being asked questions their authors couldn’t have anticipated. An internal court allows for rules to evolve in a structured manner. It creates a pressure release valve, through which existing principles are frequently reexamined, and its accumulation of precedents allows for progressive refinement and tinkering. An ambitious version might even add court opinions instantaneously to a live corpus the model can directly consult in real time.
Improving Transparency and Predictability
Right now, there is no reliable way for businesses to learn how a model will apply its vague rules to their specific use case, other than running into the failure directly and finding out the hard way. Courts can help provide this transparency, which can be a real asset for AI labs: In a competitive market where models have roughly similar capabilities, the ability to offer a more predictable business environment—one in which firms can know in advance how a model will handle their cases—is meaningful for enterprise customers dependent on stable behavior. While allowing rules to evolve over time might lead to more short-term variation, this gradual evolution is more stable in the long term than periodic top-down rewrites (in the same way judicial review reduces the need for sudden constitutional overhauls).
Giving the Public a Say
Courts could also be a promising route through which the public shapes model behavior. Amicus briefs, for example, would allow civil society groups, researchers, users, and other interested parties to explain how they think a particular case should be resolved. Most proposals for democratic input into AI specs focus on the drafting stage; involving the public on particular hard cases may be more promising. Anthropic’s Collective Constitutional AI experiment was a clear example of the difficulty of producing actionable guidelines by asking the general public to agree on broad principles from scratch. Working from narrow cases rather than abstractions makes the task much more narrow and manageable: Amicus briefs ask the public to react to a specific situation, something humans have done well for as long as there have been juries. And the option to contest at all, independent of whether contestation succeeds, is itself important.
Cultivating Public Deliberation
When the U.S. Supreme Court rules on a hard case, something happens beyond the courtroom. The opinion becomes a national argument, taught in civics classes and law school seminars. Dissents are debated at dinner tables by people who will never appear before a judge. These opinions create specific instances through which national debates about broad moral issues can crystallize and take shape. Courts, in this sense, serve a civic function analogous to what Alexis de Tocqueville saw in juries: a “gratuitous public school” encouraging men and women to engage with their reflective faculties and build habits of civic autonomy. Internal AI courts could offer something similar. The steady stream of open cases would present concrete moral questions that generate public discussion and engagement around what ought to be AI’s guiding values. Making this deliberation a “public thing,” rather than hiding it behind lab doors, could help strengthen society’s collective muscle for autonomous moral judgement.
This Is Not the Oversight Board
Eight years ago, Meta (then Facebook) launched what eventually became the Oversight Board: perhaps the closest analogy to this project, an independent court-like institution created to appeal content moderation decisions. By and large, it failed, for two key reasons.
Its deepest mistake was scale. A human court can carry out only a limited number of decisions a year, but the wisdom of the common law is to grant each appellate ruling precedential power. Though the total number of opinions is small, each has far-reaching influence. The opposite was true of the Oversight Board. Each content moderation decision was viewed in isolation, applied only to other posts containing “identical content with parallel context.” That setup allowed the board to decide very narrowly, rather than extrapolating cases into broader principles, limiting each case’s applicability and precedential force. With billions of moderation decisions a day, correcting these few instances made little difference.
Our proposal doesn’t fall prey to the same trap. Rulings on AI character are easier to implement at scale than content moderation, because the principle and reasoning behind a particular ruling can be implemented across the full range of AI behavior by training the model on natural language text. Unlike the Oversight Board, fine-tuning the character of AI models through reinforcement learning relies precisely on the ability to generalize higher-level principles with wide applicability from a few examples. This ability to abstract normative intuitions from a limited sample of examples makes character training a uniquely effective area for internal courts. In fact, recent research from Meta and Anthropic has shown that training the model on very few examples of very high quality—exactly like the in-depth reasons generated by a court opinion—allows AI to apply moral norms to very different cases.
A second key issue that plagued the Oversight Board was its paralysis problem. By designing the board to sit outside Meta, the setup created conflicting incentives, enabling it only to make recommendations that the company could deny at will. Furthermore, by creating conflicting roles that all sought decision-making power (among trustees, board members, and administrators), the setup institutionalized an incapacitating tug of war. What’s more, by making the board high profile, Meta appointed renowned board members who could only commit to working for the board part time, slowing down the process. The combination of these three factors—misaligned incentives, conflicting sources of power, and part-time staff—destined the board to be unable to move at the speed of content moderation, issuing fewer than 250 cases in total over five years, out of hundreds of billions of content moderation decisions.
But this problem isn’t inherent to private courts. Since the Oversight Board and in response to the EU’s Digital Services Act, a number of other private courts have been created to fulfill the same function. While some are plagued by the same issues, others like User Rights have operated at scale with a startup mentality: integrating cutting-edge tools, hiring aggressively, and reviewing multiple thousands of cases a year while giving each thorough consideration and detailed explanation. A dynamic and fast-moving private tech court is very much possible. Moreover, situating this court within the frontier labs also removes the adversarial and divided structure that plagued the Oversight Board. Overall, creating courts for AI constitutions, as opposed to content moderation, seems much more likely to produce rulings that scale and remain effective over time.
* * *
These decisions matter. Every day, AI models have to make millions of judgment calls to decide whether one action or another best respects their guidelines. Many of these will be easy or trivial. But plenty won’t: deciding whether a particular demand could create a biological hazard, whether a child’s question pertains to self-harm, whether a military request violates international norms. As models continue to be integrated into more and more sensitive parts of life (hospital systems, national intelligence, warfighting), these judgment calls will only become more consequential.
Yet even as these stakes rise, public debate has been stuck around what rules should bind models, paying almost no attention to how those principles are interpreted and applied. Creating courts for AI constitutions is a first real effort to shift our focus. As Daniela Amodei, president of Anthropic, once stated: “There are all of these bodies and systems that exist in the US to make sure that we follow the constitution. There are the courts, there is the Supreme Court, there is the presidency… There is all this infrastructure that you need around this document, and I feel like we are also learning that lesson here.”
More broadly, this effort fits in the broader lineage of how citizens have domesticated private sovereigns in the past. Privately held governments first took shape four centuries ago, as shareholders were put in charge of early colonial governments. Yet as these accumulated greater power over daily life, they were progressively pushed to abide by similar obligations as nation-states, including creating their own internal court systems to guarantee substantive due process and restrict arbitrary power, like the East India Company’s Mayor’s Courts. A second wave of private governance arose in 19th-century Europe and early 20th-century America with the proliferation of company towns, which made laws, resolved conflicts, and enforced punishment for everyday disputes. As these proliferated too, constitutional guarantees began to be enforced, most notably with the U.S. Supreme Court mandating that the internal rules of company towns respect free speech rights. In both instances, the substantive liberal commitments required of traditional nation-states—including courts, due process, fundamental rights, the rule of law, and so on—were deliberately extended to private Leviathans.
It is in this context that this project takes shape. After the colonial charter and company town, a third wave of private governance is now on the horizon: private general intelligence. If the leaders of these labs are even directionally correct, much of everyday life will soon be mediated by tomorrow’s frontier models. Courts for AI constitutions can help bring the substantive promises of liberalism to this new order: extending the rule of law through greater predictability, making enforcement public and transparent, giving ordinary people a way to contest specific decisions, opening public debate over which values to embed in models, and letting rules evolve over time. In doing so, these courts can lay the first stone in the much longer project of domesticating the next great private Leviathan.
