Failure Behavior

Transparency, participation, and better categories present as democratization. The question that sorts them is whether any finding can make power unable to continue.

Reading settings

Between 2016 and 2019, Australia's Centrelink raised roughly 470,000 welfare debts from a system that compared annual income records against fortnightly averages. The averaging produced numbers the actual records contradicted. When people objected that the debt described somebody they were not, the system asked for years of payslips and treated the inability to produce them as agreement.1 The averaging could not read fortnightly data. It decided anyway, and the recipient got the bill for the decision.

An institution that wants to do something consequential to a person should be able to show two things: that it is authorized to act, and that its reason actually applies in this case. When either one fails, the action stops. The administrative failures examined here run on the opposite arrangement — the authorization is presumed to persist, the reason is checked afterward if at all, and stopping is the thing the system was never built to do.

A type checker does not need to understand what your program means to reject an invalid operation. A safety interlock does not need a complete model of the factory to know that a required condition is absent. A formal system does not need to understand a person in order to know that it has failed to establish the prerequisites for acting on them. That failure can be a state the machine reports — UNKNOWN, OUT_OF_SCOPE, INSUFFICIENT_EVIDENCE, MODEL_DEFECT — and the report can stop the machine instead of the person.

A fallible institution faces conflicts among its promises. Immediate execution leaves no time for a prior challenge; irreversible action limits what later repair can recover. These are design constraints, not a single formal impossibility theorem. Deciding which promise gives way is a political choice that should be visible in the machinery.

The states a system needs

If a model is allowed to encounter cases it cannot confidently classify, either it abstains or it guesses. Forced guessing converts the model's uncertainty into action, most of it landing on whoever the guess was wrong about. Robodebt is what forced guessing looks like at population scale: the model's gap got converted into debt notices, and the people holding the gap's evidence got invoiced for the proof.

Selective-classification research studies the abstain option precisely because errors are costly, and finds reliability improves when coverage is sacrificed.2 It also finds something the reliability framing hides: the abstain option has its own distribution. A system that abstains on whom it cannot read still chooses whom it cannot read, and the choice tracks the usual measurement gap — the people the model was worst at reading are the ones least equipped to contest the reading. Someone still has to check which groups are made to wait when the model abstains.

The older default ran the other way. Institutions respond to their own uncertainty by freezing the benefit, holding the file, demanding more documentation, and letting the person pay for the wait — pending as governance. Ignorance functions as coercion: the institution doesn't know, so the person must prove. The institution's ignorance cannot by itself justify greater coercion. It can justify provisional action — an ambulance moves a person it cannot interview — but the provisional action carries its own warrant, justified on the terms of the situation rather than inferred from the gap in the file. An insurer with the same ignorance about the same person has no warrant to freeze their account, and the difference between the ambulance and the insurer is the whole argument.

UNKNOWN therefore cannot map to "no action" as a fake universal. A resuscitation bay cannot abstain; neither can a grid controller; sometimes maintaining the status quo is the harm. The constraint is: uncertainty reduces the institution's freedom to coerce, and whatever provisional burden it chooses has to be justified rather than presumed. The state exists to stop the machine from converting its own ignorance into evidence against the person.

This reframes the ambition of classification. Human systems keep answering misclassification with richer ontologies — another household type, another disability code, another eligibility exception, a better-trained model. Reality stays open-ended, every finite ontology has edges, and the edges keep arriving as cases. Unreasonable people are the edges, manufactured by the design rather than by their lives. Better classification still matters. It also needs a fallback: arrangements in which the inability to represent someone is never sufficient grounds for harming them. The universal guarantee changes with it — less that everyone fits the same boxes, more that everyone keeps standing when the boxes fail.

Diagram
Mermaid diagram
Diagram unavailable.
View source
%% title: What the failure states do
%% caption: The states replace forced guessing. A case the warrant cannot cover stops the machine and opens a challenge path, instead of deciding the person.
flowchart LR
  classDef actor stroke-width:1.6px;
  classDef system stroke-width:1.2px;
  classDef gate stroke-width:1.6px;
  classDef status stroke-dasharray:4 3;
  classDef repair stroke-width:2.4px,font-weight:bold;

  A[Case arrives]:::actor --> G{Warrant established?  
authorized + reason applies}:::gate
  G -- "yes" --> X[Act, with provenance kept]:::system
  G -- "no" --> S[UNKNOWN / OUT_OF_SCOPE /  
INSUFFICIENT_EVIDENCE / MODEL_DEFECT]:::status
  S --> Q[Stop. Justify any provisional burden]:::system
  Q --> C[Challenge reaches the rule]:::repair
  X --> C

Timing is a constitutional variable

If a person is to have a genuine opportunity to challenge the grounds of an action before an irreversible deprivation, some latency has to exist between proposal and execution. Instant execution and meaningful pre-execution contestability are available for the same irreversible action one at a time. U.S. doctrine already prices this rather than pretending it away: Eldridge balanced the private interest at stake, the risk of erroneous deprivation, and the administrative burden of additional process against Goldberg's pre-termination hearing. The Court called that due process. Most systems call it latency, and latency is a structural right whenever the action is irreversible.

That reframes where protections belong. Courts, ombudsmen, and appeals catch misuse after the fact. Some protections can move upstream, into preconditions of execution: a warrant that cannot fire unless its conditions are demonstrated, review that precedes the deprivation for the irreversible class. Later challenge still matters, and later-only challenge accepts the original architecture intact.

Contestability also consumes resources, which makes its design tri-sided: access has to be inexpensive, processing has to be bounded, and the channel has to resist capture by volume. You cannot maximize all three. An adversary with lawyers and stamina can buy the review system with volume; a fee structure that screens for seriousness also screens for poverty, sorting before proof begins. Where the triad gets resolved is visible in the design. Pretending it isn't there is how a channel stays nominal.

And there is the waiver problem. An institution can avoid examining its rule by granting exceptions. It helps one claimant and preserves the model — the exception gets filed as flexibility rather than counted as evidence about the rule. One-off settlements are how an agency buys the right to keep the rule. Mercy is real, and mercy that substitutes for correction is a pressure-release valve with the rule left standing.

Provenance or repair is a fiction

When a bolt fails on an airframe, the fleet is queryable. Every part carries a traceable history kept at execution, so one failed part grounds a fleet, finds its siblings, and gets fixed. The failure propagates into repair instead of waiting for ten thousand pilots to each discover the same defect.

Bureaucracy keeps decisions and discards dependency. Near-misses, manual overrides, the case a caseworker quietly fixed — administrative residue, discarded at the point where it would be most useful. When a rule later proves defective, the answer to "who else" may simply not exist, because the dependency was never kept and can be rebuilt by no one.

Provenance research confirms the limit: exact dependency provenance is not computable in full generality, and practical systems rely on records deliberately captured at execution time.3 Repair has to be designed where action is designed. A system that scales error automatically and scales correction manually has chosen, through its record-keeping, whom its mistakes are allowed to keep hurting.

Human review has a similar limit. A reviewer who inherits the same categories, the same evidence, and the same incentives has changed the sequence of the decision, not its grounds. The person launders the output through a conscience and passes for the in-house ethicist. Contestable-AI design work treats contestability as something designed across a system's lifecycle rather than added afterward as an explanation screen, and NIST's guidance treats human roles as design questions rather than assuming presence solves anything.4 Review needs a pathway for invalidating the governing assumption, with authority beyond overriding one output.

Self-certification is the whole collapse

If the same actor writes the mandate, defines the model, issues the warrant, adjudicates the challenges, clears the defects, and maintains the kernel, the constraint system has compiled to one sentence: Agency A may act whenever Agency A says Agency A may act. Nothing visible contradicts it, because everything visible is Agency A.

The 2008 rating agencies ran that sentence in public. The issuer paid the validator; the same institutions defined the models, validated the securities, and fielded the doubts about their own validation, and triple-A came to mean what the issuer needed it to mean. Self-certification left the machinery looking rigorous the whole way down.

So one structural result survives: a revocable-authority architecture requires at least one locus of validation the exercising principal cannot unilaterally rewrite. That is separation of powers as a dependency constraint. It does not say how many branches, or who appoints them. It says the number cannot be one, and the principal cannot sit as its own validator.

The boundary stays two-sided. The subject's power is the power to compel revalidation; confirmed failure removes authority; unconfirmed challenge changes nothing about a rule's authority to proceed while its warrant is checked. Otherwise the architecture swaps administrative sovereignty for claimant sovereignty, and nothing serves anyone at scale. And some legal rules are mechanical while others carry interpretation, discretion, and judgment no schema exhausts. The OECD's law-as-code consultation keeps interpretive elements as such and leaves the authoritative text legally controlling.5 You can compile limits on power without compiling away judgment. Deciding what warrants ought to exist is a different problem from enforcing them, and politics stays above the constraints — the machinery's job is honest enforcement, not legislation.

The sorter

So the sorting question for anything presenting as democratizing: does the intervention change when power is allowed to proceed, who can force it to justify itself, and what happens when justification fails? If the answer is no, it may still be useful, and it is still not a solution to domination. Several families fail the test predictably.

Transparency without revocability: audit logs, model cards, published reasons — and the institution keeps acting after everyone can see the justification is bad. Legible domination is still domination.

Participation without interrupt power: listening sessions, advisory councils, stakeholder engagement — channels that absorb dissent while discretion stays put. The ladder already sorted voice from corrective standing; everything below corrective standing is consultation.

Technocratic humility without authority loss: "we know models are imperfect, we monitor bias, we continuously improve." Acknowledged uncertainty with no effect on executability is apology with a dashboard.

Decentralization as a substitute for corrigibility: moving power closer can help, and a local institution can dominate as effectively as a national one. Decentralization can also become the excuse for abandoning universal provision, leaving employers, landlords, and markets uncontested in the vacuum. Small is not synonymous with non-dominating.

Representation alone: more kinds of people inside the institution improves knowledge and shifts priorities, and if insiders still hold final authority over whether their own categories have failed, representation changes who operates the sovereignty. The affected person still cannot compel a revision.

Purity politics deserves its own line. A politics organized around whether a person or statement belongs to the correct moral category has rebuilt the ontology problem inside the movement: people turn into instances, ambiguous cases turn suspect, disagreement reads as contamination. A movement can get extremely good at classification while losing the capacity for correction. The category problem is internal to it now.

Responsible AI is the contemporary test case. Audits, model cards, ethics boards, red teams, human review — each useful; none answers what finding actually makes the system unable to continue doing the thing. If the answer is "whatever management decides," the governance layer is advisory.

Reforms sort by what they touch. Making power more informed, transparent, representative, careful, or benevolent touches the rulers. Making power unable to proceed when its justification fails touches the failure behavior. Neither good rulers nor good rules are enough: emancipatory politics has to change the failure behavior of power itself.

Permission that must renew

We already bound software agents this way: scopes, revocable credentials, audit trails, dependency graphs, the capability to act bounded by explicit warrants. The same architecture applies upward. A hospital, an insurer, a benefits agency, a platform can be capable, large, automated — and stoppable. Capacity and humility stop being opposites once authority has to renew instead of persist.

And the people most exposed to institutional error are part of the mechanism by which the institution discovers the limits of its own authority. Their objections can reveal defects the institution would otherwise keep applying.

When the reasons for power fail, the power should stop — not the person. That requires someone who can suspend the action, records that identify others affected, and protection for the people waiting for repair.

Notes

1

Royal Commission into the Robodebt Scheme, Final Report (Commonwealth of Australia, 2023). The Commission found the debt-raising scheme unlawful: averaged annual data stood in for fortnightly records, and the onus of disproving sat with recipients. See robodebt-royal-commission.gov.au.

2

Selective Classification Can Magnify Disparities Across Groups (arXiv:2010.14134). Abstention improves aggregate reliability and can shift its burden onto already-disadvantaged groups — the abstain state is necessary and has its own politics.

3

Provenance as Dependency Analysis (arXiv:0708.2173). Exact dependency provenance is intractable in general; provenance capture is a design decision made at execution, not a reconstruction available afterward.

4

Zaga, de Ruiter, Houben, and Liao, "Contestable AI by Design: Towards a Framework," Minds and Machines (2022), link.springer.com; NIST, AI Risk Management Framework, Appendix C on human-AI interaction, airc.nist.gov.

5

OECD, Consultation on the digital provision of law: Towards a shared reference framework for law as code (2026), oecd.org.

Continue reading

Related essays

Next routes

Continue with the next essay in Conditional authority, or widen into Core claims of the archive once the sequence clicks.

Version history

No prior versions in this archive snapshot.

    Get new essays by email

    An occasional note when a new essay goes live.

    Get new essays by email

    An occasional note when a new essay goes live.