Take a subscription company whose billing service charges some customers twice for a month. A deploy changed how retries were handled, a payment processor timed out on a busy afternoon, and for a few hours the retry logic treated a slow success as a failure and charged again.
The engineering response is exemplary. Someone spots the anomaly in a dashboard within a day. The team writes a postmortem that names no culprit, lays out the timeline, and traces the defect to a missing idempotency check and a review process that had no way to catch it. There are action items with owners and due dates. The fix ships with a test. The engineer who merged the change is thanked for writing half the document. Nobody is fired, because the organization has learned that firing the person who happened to be holding the pen teaches it nothing about the pen.
Then the incident is closed. The double charges are still sitting on thousands of card statements. Customers who notice and write in get a refund, usually quickly and with an apology. Customers who don’t notice keep the loss. Nobody went looking for them, although the same logs that told the engineers which requests were retried also say, to the row, which accounts were charged twice.
One incident, two ledgers
The blameless postmortem is one of the better things modern organizations have learned to do. Google’s site reliability engineers made it a published practice: a written account of an incident that assumes the people involved acted reasonably on what they knew, and asks what about the system let a reasonable action cause harm.1 John Allspaw’s account from Etsy made the same case, and Sidney Dekker’s work on just culture gave it a longer history in aviation and medicine.2 The logic is sound. People make errors because systems allow them to, and punishing the person leaves the system ready to produce the same error through someone else.
That logic has an internal and an external side, and organizations apply only the first. Inside, the incident is treated as a property of the system. Nobody has to prove they were not negligent; the organization carries the investigation and the fix. Outside, the incident is treated as a matter for each customer. The customer has to notice the charge, recognize it as an error rather than a price change, find the support channel, describe the problem in terms the agent will accept, and wait.
So the asymmetry is precise. An engineer is protected from absorbing the failure personally. A customer is required to absorb it personally, until they hand it back, so blamelessness stops at the payroll.
The archive has argued that responsibility should land where authority lived: the power to have altered the outcome, when it was still alterable.3 A blameless postmortem honors that rule inside the building. It moves the fault off the engineer and onto the design, which the organization owns. But an organization that has accepted, in writing, that its design caused the harm has also accepted the harm. Owning the cause and disowning the consequence is a strange combination, and it is the ordinary one.
The same essay warned that the alternative to scapegoating “is not a blameless institution,” which it called “the fantasy of every system that prefers its errors unexamined.” The two positions fit once the person and the organization are held apart. A blameless postmortem lifts blame off an individual and sets it on the design, and the design belongs to the organization. Blamelessness for people is the rule that essay argued for. Blamelessness for the organization is the fantasy it named, and a claims process is how an organization that has cleared its staff arrives there: the harm its design caused becomes nobody’s, and whoever it landed on is left to prove it happened.
Why the line runs along the payroll
Part of the answer is bookkeeping. An incident, in most reliability practice, is a disruption to a service. It opens when the service degrades and closes when the service is restored and the follow-up work is tracked. Refunds belong to a different department with a different budget and a different measure of success. Engineering is judged on recurrence. Support is judged on handle time and satisfaction scores. Finance sees refunds as a cost to be forecast and, where possible, contained. No one in that arrangement owns the number of people who were harmed and have not been made whole, so no one reports it.
The other part is that remediation is designed as a claim. A refund that has to be requested is a refund most people will never receive, and the organization knows this without having to decide it. The claim rate is a variable, and every point it falls below one hundred percent is money that stays on the company’s books. Nobody needs to intend that outcome for the arrangement to produce it reliably. The sorting happens in the gap between an error that occurred and an error that was noticed, and the customer is the only party assigned to close that gap.
Noticing is labor, and it is unevenly distributed. The customer who reconciles every statement catches the double charge. The customer working two jobs, or caring for a parent, or whose card statement goes to an email address they check once a month, does not. A patient billed for a procedure that was covered, a claimant whose benefit was miscalculated, a tenant charged a fee the lease didn’t allow: in each case the institution’s error is corrected in proportion to the victim’s spare attention. The people with the least of it pay the most for the institution’s mistakes, which is the same routing the archive keeps finding wherever software runs cleanly because a person is soaking up the variance.4
The claim process also changes what the error looks like from inside. A refund total drawn only from complainants is smaller than the harm, and when the organization later asks how bad the incident was, it reads that total and concludes: contained. The customers who never noticed are absent from every report, and a harm that appears in no report cannot shape the next decision about how much to spend preventing it.5
Carrying the logic outward
Extending blamelessness past the payroll means applying the postmortem’s own premises to the people the failure reached. A customer who missed the double charge was doing what ordinary people do with a card statement. A system that relies on their noticing is a system that depends on exceptional conduct from people it has no right to conscript.6 If the organization would not accept “the engineer should have been more careful” as a root cause, it has no standing to accept “the customer should have checked their statement” as a remediation plan.
In practice, proactive remediation has a few parts, and none of them is exotic.
The organization identifies everyone affected, using the same data it used to diagnose the defect. In most billing, benefits and claims failures this is a query, and the engineers have usually already written it to size the incident.
It makes them whole without a claim: the refund goes back to the original payment method, the benefit is recalculated and paid, the wrong denial is reversed and the patient told. Interest or a fee reversal follows where the money was held long enough to cost the person something.
And the incident stays open until the external harm is repaired. The postmortem template gets one more required section: the number of people affected, the number made whole, and the date the last one was paid. An incident with unrepaired victims is an incident in progress, and it should sit on the same dashboards and in the same review meetings as an unfixed defect.
A stopped Toyota line could not restart until the problem was fixed.7 The equivalent rule for remediation is that the organization cannot declare the matter over while it is still holding what the defect took.
Some regulators already require something like this in particular sectors, directing firms to identify affected customers and pay redress without waiting for complaints.8 The principle underneath is general, and an organization can adopt it without waiting for a regulator, provided someone other than the organization checks the count. A make-whole the company administers and declares finished is the claim process under a kinder name.
The objections
The first is cost, and it is correct: finding and repaying every affected person costs more than repaying the ones who complain. The difference is exactly the amount the organization has been keeping. An argument that proactive remediation is too expensive is an argument that the error was profitable.
The second is that some affected people can’t be identified or reached. Sometimes true, and the right response is a documented effort and an unclaimed balance that stays owed, rather than a quiet write-back to revenue. The existence of a hard residue is no reason to skip the easy majority.
The third is fraud: an open offer invites false claims. Proactive remediation needs no offer. The organization pays the accounts its own logs identify, and the logs are the one witness nobody can coach.
Closing the incident
A blameless postmortem is a promise that the organization will treat its failures as its own. Most organizations keep that promise to their staff and break it with everyone else, and the break is visible in a single field. Look at when the incident was marked resolved and then at when the last affected customer got their money back. In most organizations the second date doesn’t exist, because nobody was assigned to produce it, and the customers who never noticed are still paying for a defect the company fixed, documented and learned from.
Notes
Betsy Beyer, Chris Jones, Jennifer Petoff and Niall Richard Murphy, eds., Site Reliability Engineering: How Google Runs Production Systems (O’Reilly, 2016), chapter 15, “Postmortem Culture: Learning from Failure.”
John Allspaw, “Blameless PostMortems and a Just Culture,” Code as Craft (Etsy engineering blog), May 2012; Sidney Dekker, Just Culture: Balancing Safety and Accountability (Ashgate, 2007).
Responsibility is legitimate only where the authority lived: The Nurse Has No Committee.
The system stays smooth while the person scrambles: Everyone is a Crumple Zone Now.
Who else was hit, when nobody kept the record: Stop the Machine, Not the Person; and procedures that turn complaint into attrition: Nothing Turns Into a Decision.
Correction that depends on the vigilance of the affected is a burden transfer with a grievance procedure attached: The Right to Get Tired.
Stopping power, and a stopped line that restarts only once the problem is fixed: The Stop Button Was Removed.
The UK Financial Conduct Authority’s Consumer Duty (PRIN 2A, in force from July 2023) requires a firm that finds retail customers have suffered foreseeable harm from its acts or omissions to take appropriate action to rectify it, including redress where appropriate, whether or not anyone has complained.