AI alignment has been allowed to expand far beyond the problem it was built to solve. It began as a technical question, how to make an artificial agent reliably pursue intended ends, and it has taken on the prestige of a theory of how consequential technical power should be governed, which it is not.
Alignment is primarily a theory of successful delegation. Governance is a theory of legitimate authority. A principal asks whether an agent will do what the principal intends. The person the system acts on has to ask something more basic. Why does this institution have authority over me, what evidence may it use, can I inspect and correct that evidence, what happens while I challenge its decision, who can actually reverse the outcome, and who bears the cost when the institution is wrong?1 Making the machine more obedient does not answer those questions.
I am not claiming that alignment researchers believe alignment covers all of governance. Nick Bostrom, OpenAI, Anthropic, and contemporary technical researchers say nothing so crude. OpenAI's 2023 Superalignment announcement acknowledged that managing superhuman systems would require new governance institutions alongside technical control.2 Anthropic's Collective Constitutional AI research notes developer power, selection bias, and unresolved questions of democratic legitimacy.3
The real problem is harder to dismiss. Alignment keeps supplying the conceptual grammar in which broader questions of AI power get framed, and that grammar favors the relationship between principal and agent while leaving the relationship between institution and subject underspecified. It mistakes one axis of governability, the faithfulness of delegated power, for the fundamental one. That first substitution sets off a chain of others:
- preference learning stands in for legitimacy
- oversight stands in for accountability
- human control stands in for democratic control
- corrigibility to principals stands in for corrigibility to subjects
- value aggregation stands in for political settlement
- obedience stands in for safety
My counterclaim is this.
The fact that power faithfully serves an authorized objective tells us almost nothing about whether that objective may legitimately be imposed on the person who bears its consequences.
Constitutionalism starts from the position of the person on the receiving end, where alignment starts from the principal's chair. Seen from there, much of what is treated as new in AI governance has precedents. Administrative law already knew that benevolent purpose is not due process, and labor history knew that apparent automation can rest on invisible human adaptation. The problem of governing abstraction predates AI, and what AI changes is the scale, speed, and reach at which old forms of abstraction can be put to work.
The residual
Institutions cannot act on reality without simplifying it. The patient becomes a record, the household an eligibility unit, the neighborhood a risk profile. Some compression is unavoidable because organized action requires representations. The political problem begins when the representation is treated as a complete account of what it describes.
James C. Scott showed how administrative legibility makes populations governable.4 Geoffrey Bowker and Susan Leigh Star showed how classification systems lock choices about knowledge and politics into infrastructure.5 The question here is what happens to what the representation leaves out. Every formal system produces a residual, the part of reality it cannot represent but still has to handle in order to work, and someone downstream carries it. An automated eligibility system returns a decision in milliseconds while the misclassifications, missing documents and failed validations go through an unrecorded channel of friction. The institution records the milliseconds and the claimant absorbs the weeks. Leaving that labor out of the accounting lets an institution look efficient while it moves the cost of fixing its errors onto the people it governs.6
Alignment as principal-agent optimization
The central limitation of alignment follows from its starting point. It treats getting the objective, the preference relation, the constitution, or the supervisory structure right as doing far more political work than it can. Consider the strongest formulations the literature has produced.
Stuart Russell's framework of "provably beneficial AI" replaces fixed objectives with machines that are uncertain about human preferences and learn them from human behavior.7 The Cooperative Inverse Reinforcement Learning literature formalizes value alignment as a cooperative game in which the robot and the human share the human's reward function.8 The constitutional response is immediate. Beneficial to whom, acting under whose authority, against whom, and with what rights of refusal and correction?
Accurately inferred human welfare does not imply legitimate authority to optimize it.
Suppose an AI system knows with perfect accuracy what a population would prefer in aggregate. It still does not follow that a hospital, insurer, employer, agency, or platform is entitled to impose the resulting decision on a particular person. Russell's proof is about the relationship between action and preferences. It says nothing about standing, jurisdiction, rights, procedure, or remedy.
Eliezer Yudkowsky's Coherent Extrapolated Volition poses the normative problem as discovering what humanity would want if we "knew more, thought faster, were more the people we wished we were," and converged where our volitions cohered.9 The usual objection is that human preferences do not converge. The deeper one asks why convergence, even if it existed, would give anyone authority. Suppose CEV succeeded completely. You would still have to ask who is entitled to act in humanity's name, what cannot be done to dissenters, and which domains should not be governed through preference maximization at all. Some political constraints exist precisely because knowing what most people want does not settle what may be done.
Anthropic's Collective Constitutional AI experiment asked roughly 1,000 Americans to deliberate on principles that went into a model constitution.10 Anthropic is candid about its limits. Yet the experiment treats gathering preferences as the main problem in democratizing AI.
Democratic preference input is not the same as democratic control over power.
Imagine that a flawless deliberative process produces a system's constitution. You still have not answered who may deploy it to determine welfare eligibility, whether an affected person can inspect and challenge its inferences, whether an appeal suspends the deprivation, who is liable for a catastrophic false positive, or whether certain decisions may be automated at all.
The same limit applies to treating political disagreement as utility aggregation or mechanism design. A criminal defendant does not need their preference about conviction entered into a social welfare function. What they need are rights, vetoes, jurisdictional limits, evidentiary burdens, and standing to contest another actor's exercise of power. Sometimes the problem is not that the sovereign has misunderstood what people want. Sometimes the sovereign is not entitled to decide.
Oversight versus standing
The vocabulary of "human oversight" hides this distinction because it can describe two different relationships, managerial auditing and adversarial standing. OpenAI describes scalable oversight as using AI systems to help humans evaluate other AI systems on tasks too hard for humans to judge directly.11 The mistake is the political substitution that follows.
"Someone supervises the AI" is read as if it meant "the exercise of power is accountable".
It does not mean that, because the supervisor works for the institution exercising the power and the person the decision lands on may have no relationship to that supervisor at all. An agency may employ conscientious overseers who audit a fraud-detection model continuously, and the person hit by the model's error still needs something else. "I can force you to examine my records, inspect your evidence, halt the termination of my income, and reverse your decision." That is a claim of adversarial standing, and managerial supervision cannot supply it. An overseer evaluates a system to improve it, and a subject with rights challenges an institution to protect themselves. A preference signal is not a hearing, and red teaming is not an appeal.
%% title: Managerial oversight vs adversarial standing
%% caption: Alignment models managerial supervision of an agent by a principal. Constitutional governance requires adversarial standing for the subject against the institution.
flowchart TD
classDef actor stroke-width:1.6px;
classDef system stroke-width:1.2px;
classDef gate stroke-width:1.6px;
classDef binding stroke-width:3px,font-weight:bold;
classDef status stroke-dasharray:4 3;
classDef burden stroke-width:2px;
classDef repair stroke-width:2.4px,font-weight:bold;
classDef failure stroke-width:2.2px,stroke-dasharray:6 3;
subgraph Alignment["Alignment Paradigm: Managerial Oversight (Principal -> Agent)"]
direction TB
P1["Principal / Auditor"]:::system --> O1["Scalable oversight & red teaming"]:::system
O1 --> M1["Objective alignment & model cards"]:::status
M1 --> G1(["Systemic obedience & performance"]):::repair
end
subgraph Governance["Governance Paradigm: Adversarial Standing (Subject -> Institution)"]
direction TB
S1([Governed Subject]):::actor --> G2{"Enforceable standing"}:::gate
G2 == "legal discovery & notice" ==> A1[["Halt termination & inspect records"]]:::binding
G2 == "due process" ==> A2[["Compel reconsideration & reversal"]]:::repair
G2 -. "if denied standing" .-> D1([Absorbs error as moral crumple zone]):::failure
end
Alignment ~~~ Governance
The same gap opens when Russell's notion of deference enters an institution. Uncertainty over human preferences, Russell observes, produces useful hesitation. Machines will ask permission, defer to human guidance, and allow themselves to be switched off.12 With a single human in a toy environment, this is an elegant control mechanism. Inside an institution it hits a question it cannot answer. Who is "the human"? Is it the caseworker or the agency director, the insurer or the patient, the drone pilot or the civilian inside the target perimeter? Once several humans hold unequal positions, "defer to humans" stops specifying anything, because it has run into the problem of power and jurisdiction.
So "human in the loop" is a shell game unless it says what the human can actually do. Madeleine Clare Elish's concept of the moral crumple zone captures part of this: responsibility collapses onto the operator nearest the failure even after meaningful control has moved elsewhere.13 The pathology is splitting burden from authority, and leaving humans out of the loop is only one way to do it.
Administrative law got there first
Administrative law has been working on notice, evidentiary burden, standing, suspension, appeal, and remedy for generations. In Goldberg v. Kelly the Supreme Court held that welfare could not be terminated without a hearing first, because a recipient deprived of subsistence cannot survive the wait for a later appeal, and Mathews v. Eldridge turned that into a test weighing the private interest, the risk of erroneous deprivation, and the value of added safeguards against the government's burden.14 The American administrative state is no model of legitimacy, but it had learned that accuracy is not the whole problem of authority and that timing, procedure, and who carries the cost of error matter as much. Alignment did not discover governance. It arrived inside a much older argument about the legitimate exercise of fallible power.
Whose corrigibility
An automated welfare bureaucracy can be exquisitely corrigible to agency leadership and not corrigible at all to the claimant it wrongly cuts off. Alignment corrigibility means the ruler can correct the instrument, and political corrigibility means the people being ruled can force corrections on the ruler and the instrument together.15
The nightmare of the subject
Many rights are deliberately inefficient. Appeals duplicate work and independent review delays decisions. They exist because societies learned that smooth execution becomes dangerous when joined to coercive authority. A procedural veto is a constitutional safeguard, even if an operations dashboard would log it as friction. Other freedoms rested on nothing more than what institutions could not yet afford to do, and cheap capability removes those faster than anyone replaces them with rights.16
The existential-risk worldview centers on an AI that disobeys legitimate human intentions.17 The equal and opposite danger is the system that perfectly obeys illegitimate authority.
Misalignment is the nightmare of the principal. Perfect alignment can be the nightmare of the subject.
An obedient benefits sanctioner is not an alignment failure. Its danger is how flawlessly it carries out what the institution intends. "Make the system reliably do what its authorized principal intends" cannot be a theory of safety, because sometimes the authorization structure is the danger.
Bilateral accounting
The politics of the residual comes down to a principle of bilateral accounting. When an institution simplifies reality for its own operational advantage, it should by default bear the cost when the simplification fails. The institution that chooses the category should investigate the exception, and one that automates the judgment should finance the appeal. If the model cannot represent the person, the person should not have to reorganize reality until they become legible to it.18
This reverses the usual arrangement, in which convenience flows upward and residual complexity flows downward, and nothing about that arrangement is technologically inevitable. Someone decides whether appeals count as operating cost, whether recurring exceptions indicate a broken category or defective people, and whether ninety-nine percent accuracy counts as success without asking what happens to the remaining one percent. These are political decisions made in the language of system design.
Authority should not exceed corrigibility
The real alternative to alignment doctrine is a full theory of technical authority, in which alignment has a necessary but subordinate role. That theory asks whether the system behaves as intended. It also asks whether the institution has legitimate jurisdiction, who has standing to challenge a decision, whether action can be suspended while it is challenged, and what remedy follows when the system is wrong.19
For consequential systems, enforceable safeguards are part of what it means to be correct. A decision that cannot be reconstructed is incomplete. An override that cannot safely be exercised is fictitious. An appeal that arrives only after irreversible harm is not meaningful correction. A system that learns from mistakes while leaving the people harmed by them uncompensated is improving itself at somebody else's expense, and one that keeps strong aggregate performance by concentrating error on the people least able to challenge it is politically organized, not merely inaccurate.
Authority should not exceed corrigibility.
The more power an institution gains through technical systems, the more it has to protect the ability of affected people to inspect, interrupt, challenge, and reverse that power, and to get a remedy from it. Technocratic reasoning counts the decision but not the appeal, and the prediction but not the correction burden imposed when the prediction is wrong. A more rigorous account keeps those costs inside the ledger.20
Better models will not eliminate the residual, because every representation leaves something unresolved. Calling what is left an edge case does not give it a custodian, and the fight is over who that custodian is, what power they get in return, and whether the institution that created the residual has to pay for correcting it.
Here alignment runs out of explanations. Beyond it lies the older and harder question: under what conditions does power deserve to be exercised at all?
Notes
The distinction between what a category drops and who carries it was first introduced in Nine Thousand Claims and One Woman. The physics of residual stress as an unabsorbed load () was formulated in Politics Is Applied Physics.
Jan Leike and Ilya Sutskever, "Introducing Superalignment" (OpenAI, 2023).
Anthropic, "Collective Constitutional AI: Aligning a Language Model with Public Input" (2023).
James C. Scott, Seeing Like a State: How Certain Schemes to Improve the Human Condition Have Failed (Yale University Press, 1998). Scott asks what gets destroyed when states simplify reality. This essay asks who must compensate for the simplification after it is imposed.
Geoffrey C. Bowker and Susan Leigh Star, Sorting Things Out: Classification and Its Consequences (MIT Press, 1999).
Where the accounting boundary is drawn, and how leaving out the repair labor makes displacement look like efficiency: Per What? The Denominator Is the Doctrine.
Stuart Russell, Human Compatible: Artificial Intelligence and the Problem of Control (Viking, 2019).
Dylan Hadfield-Menell, Stuart J. Russell, Pieter Abbeel, and Anca Dragan, "Cooperative Inverse Reinforcement Learning," Advances in Neural Information Processing Systems 29 (NeurIPS, 2016).
Eliezer Yudkowsky, "Coherent Extrapolated Volition" (Machine Intelligence Research Institute, 2004).
Anthropic, "Collective Constitutional AI" (2023), cited above.
OpenAI, "Our Approach to Alignment Research" (2022), and Leike and Sutskever, "Introducing Superalignment" (2023), which names scalable oversight; see also Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei, "Deep Reinforcement Learning from Human Preferences," NeurIPS (2017).
Russell, Human Compatible (2019), cited above.
Madeleine Clare Elish, "Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction," Engaging Science, Technology, and Society 5 (2019). See also Everyone is a Crumple Zone Now.
Goldberg v. Kelly, 397 U.S. 254 (1970); Mathews v. Eldridge, 424 U.S. 319 (1976). What the timing argument demands of an automated system: Stop the Machine, Not the Person and The Rule Is Never on Trial.
Nate Soares, Benya Fallenstein, Eliezer Yudkowsky, and Stuart Armstrong, "Corrigibility," AAAI Workshop on AI and Ethics (2015). The inversion from principal to subject, worked through: The Corrigible Machine; why transparency without a binding return channel fails: All the Machinery, None of the Correction.
What cheap capability does to limits that were only ever prices: Nobody Exceeded Their Grant.
Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014); Eliezer Yudkowsky, "Creating Friendly AI" (Singularity Institute, 2001).
On adaptation privilege: Rank Is the Right to Remain Rigid.
The machine states (UNKNOWN, OUT_OF_SCOPE) and the requirement for revocable authority: Stop the Machine, Not the Person.
For the visual and operational model of bilateral accounting, see The residual ledger.