Apply the framework here
In AI governance, the framework’s claim becomes: oversight fails legitimacy when humans remain in the loop ceremonially while contestation remains costly and weak.
The question is not whether principles exist; it is whether harmed people can trigger correction fast enough to matter.
That capacity sits in the institution rather than the model. One system can be survivable in the first hospital that installs it and dangerous in the second, and what separates them is staffing, escalation authority, and whether a bad number obliges anyone to act.
Recognition
Common misdescription in this field
Misdescriptions common in AI oversight discourse.
Structural decisions are laundered as objective predictions.
- Governance language obscures category politics.
- Error is treated as residual noise.
- Affected people inherit correction burden.
Human sign-off is used for legitimacy while power sits in model defaults and organizational tempo.
- Operators hold liability without design authority.
- Escalation channels are narrow and slow.
- Override exists in policy but not in practice.
Evaluation scores travel between institutions; the conditions that make an error survivable do not.
- Accuracy is measured where nobody has to live with the residue.
- The same model meets a different staffing ratio in every building.
- Risk is a property of the pairing, not of the system being installed.
A burden measure improves fastest when the work moves outside whatever the measure counts.
- Correction passes to contractors, temps, or the user.
- Categories narrow until the expensive cases stop qualifying.
- Reported volume falls while unreported rework continues.
Operational diagnostics
What to measure instead
Measure contestability and repair, not policy prose — and measure the institution receiving the system, not only the system.
Deployment substrate: what happens when this is wrong here?
Run survivability, reversibility, burden, standing, and time against the receiving institution before the model arrives. The spread between two buildings running one system is the part no model card reports.
Appeal burden: who can realistically contest an output?
Track time, evidence load, expertise requirements, and retaliation risk.
Override efficacy: can a human actually alter outcomes?
Measure real override rates and time-to-correction.
Remedy load: who does a granted appeal land on?
Every appeal right is caseload for whoever answers it. Recourse issued without staffing is burden transfer to the reviewer, and it surfaces as backlog rather than as a refused right.
Trigger: which number, at what level, compels which action?
A published metric with no pre-committed consequence is a report. Name the threshold, the action it forces — investigate, throttle, withdraw — and the office obliged to take it, before collection starts.
Failure dynamics
Typical failure pathway (how people fall out)
Typical AI-governance failure pathway.
Interventions
Design/legal/operational fixes
Fixes must make recourse enforceable, fast, and costly to fake.
Embed user-visible explanation, appeal, and rollback in product flows.
Guarantee timely human correction, fund the desk that has to deliver it, and penalize noncompliance.
Publish appeal volumes, reversals, latency, and unresolved cases.
Bind appeal rate, reversal rate, and time-to-correction to a named action agreed before launch, so a bad number starts something without anyone volunteering.
Scope the count to wherever correction work lands, including contractors, other departments, and people outside the organization.