Skip to content
AI.info

Ethics & Governance

Algorithmic Accountability: Who Is Responsible When AI Fails?

When AI systems make mistakes that harm people, who is accountable? Exploring liability frameworks, corporate responsibility, and emerging legal standards for AI accountability.

Algorithmic Accountability: Who Is Responsible When AI Fails?

Gabriele Masetti ·

The Excuse That Should Have Died a Decade Ago

"The algorithm made the decision" is not an explanation. It is a laundering operation. It takes a choice made by a company that built a system, a government agency that bought it, and a set of managers who deployed it without adequate testing, and it relocates the blame onto a piece of software that cannot be sued, cannot be fired, and cannot go to prison.

A decade of documented failures — in criminal justice, welfare administration, hiring, healthcare, and policing — shows exactly where responsibility actually sits: with the institutions that built, bought, and deployed these systems without the safeguards they knew, or should have known, were necessary. The evidence no longer supports treating algorithmic harm as a diffuse, no-fault event. It supports assigning liability the way we assign it everywhere else human institutions cause harm: to the party that had control and the party that profited.

The starting case: COMPAS and the diffusion of blame

In May 2016, ProPublica journalists Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner published "Machine Bias," an investigation of Northpointe's COMPAS recidivism-risk tool, used by courts across the United States to help set bail and sentencing decisions. Their analysis of over 7,000 defendants in Broward County, Florida, found that Black defendants who did not go on to reoffend were rated high-risk at nearly twice the rate of white defendants who did not reoffend — a 1.9-times-higher false-positive rate.

Northpointe disputed the framing, arguing that the tool was equally accurate in its predictive value across race and that ProPublica had chosen the wrong fairness metric. That academic dispute — separation versus sufficiency, in the technical literature — was real and unresolved. But it obscured a simpler fact: a proprietary tool was influencing decisions about human liberty in courtrooms across the country, and neither the defendants, nor often the judges, could inspect how it worked.

Northpointe treated its scoring formula as a trade secret. Courts treated the output as sufficiently authoritative to act on. When someone asked "who is accountable for this score," the honest answer was: no one, by design.

Case Harm scale Outcome
COMPAS (ProPublica, 2016) 1.9x higher false-positive rate (7,000+ defendants analyzed) Disputed methodology
Dutch childcare benefits (2005–2019) 26,000–35,000 families wrongly accused Government resigned Jan 2021
Michigan MiDAS (2013–2015) ~40,000 wrongly accused, up to 93% error rate $20M settlement (2022), ~3,000 covered

That design pattern — proprietary weighting, opaque inputs, institutional reliance — is the through-line connecting every case that follows. The technology changes; the accountability vacuum does not.

The Netherlands shows what happens when a state hides behind the vacuum

The clearest case for assigning responsibility to institutions rather than to "the algorithm" is the Dutch childcare benefits scandal, or toeslagenaffaire. Between 2005 and 2019, the Dutch tax authority, the Belastingdienst, used a self-learning risk-classification algorithm to flag childcare-benefit applications for fraud. The system used dual nationality and non-Dutch-sounding names as risk factors, effectively encoding ethnic profiling into an official government function.

At least 26,000 parents — some estimates run to 35,000 — were wrongly branded as fraudsters and ordered to repay benefits in full, often tens of thousands of euros, with no meaningful appeal process. Families lost homes, went bankrupt, and in more than 1,600 cases had children removed by youth-protection services. The scandal brought down the government: Prime Minister Mark Rutte's third cabinet resigned on January 15, 2021, less than two months before a general election.

This is the model case because the chain of responsibility is not actually complicated once you refuse the algorithmic excuse. A government agency chose which variables to feed a model. It chose to treat model output as grounds for aggressive, irreversible enforcement action. It chose not to build in human review with the power to override, and when caseworkers did flag problems, the system's institutional momentum overrode them.

No independent audit function existed to catch the discriminatory proxy variables before tens of thousands of families were harmed. Every one of those choices was made by identifiable humans inside an identifiable institution. The algorithm did not resign the government; the government's decision to deploy an unaudited, ethnically discriminatory system without recourse did.

Private-sector deployment carries the same logic — Amazon and Optum

Critics of assigning hard liability sometimes argue that private companies deserve more latitude because market discipline will correct their errors faster than government bureaucracy. Two of the most-cited private-sector cases undercut that argument.

Amazon spent from 2014 until 2017 developing an internal hiring tool to rank job applicants' resumes on a one-to-five star scale. Reuters reported in October 2018 that Amazon scrapped the project after discovering it had taught itself to penalize resumes containing the word "women's" — as in "women's chess club captain" — and to downgrade graduates of two all-women's colleges, because it had been trained on ten years of resumes submitted to a company whose technical hires skewed heavily male.

Amazon's engineers tried to patch the specific terms the model had learned to penalize, then abandoned the project because they could not be confident the system would not find other, subtler proxies for gender. That is, in fact, a reasonably responsible outcome — the company caught the problem internally before the tool was ever used to make real hiring decisions, and killed it rather than ship it.

The case is instructive precisely because it shows the alternative to accountability failure: internal testing, willingness to kill a product line, and transparency about why. Compare that to Northpointe's continued defense of COMPAS or the Belastingdienst's yearslong denial, and the difference is entirely institutional will, not technical inevitability.

The Optum case, published by Ziad Obermeyer and coauthors in Science in October 2019, is the harder problem, because the discrimination was not an obvious slur-proxy like "women's chess club" — it was a subtle, structural choice buried in the model's design. A risk-prediction algorithm used across U.S. hospitals to identify patients for high-risk care-management programs used prior healthcare spending as a proxy for healthcare need.

Because less money is spent, on average, treating Black patients with the same level of underlying illness as white patients — a function of unequal access, not unequal need — the algorithm systematically scored Black patients as healthier than equally sick white patients. The researchers estimated this proxy error reduced the number of Black patients identified for extra care by more than half: correcting it would have raised the share of Black patients flagged for additional help from 17.7 percent to 46.5 percent.

Retraining the model on a combination of cost and clinical variables such as chronic conditions cut the racial disparity by 84 percent. The fix was neither exotic nor expensive. What was missing, before Obermeyer's team intervened, was any institutional obligation to audit the tool's outcomes by race before or during deployment. The hospitals and insurers using it had accepted a vendor's design choice — cost as a proxy for need — without independently verifying whether that choice was safe. That is a governance failure, not a math failure.

Obermeyer et al., Science (2019): correcting a cost-based proxy for health need sharply changed who got flagged.

When the false positive is a human being in handcuffs

Facial recognition turns the same governance failure into physical harm. Joy Buolamwini and Timnit Gebru's 2018 "Gender Shades" study tested commercial gender-classification systems from IBM, Microsoft, and Face++ and found error rates under 1 percent for lighter-skinned men but as high as 34.7 percent for darker-skinned women — a roughly 35-fold gap hiding inside tools marketed as broadly accurate. That single number reframed how regulators and the public understood facial-recognition risk: aggregate accuracy figures were concealing a system that functioned far worse for exactly the population most likely to be over-policed.

The consequences of ignoring that finding arrived two years later. In January 2020, Detroit police wrongfully arrested Robert Williams in front of his wife and daughters after a facial-recognition search matched a blurry surveillance still from a shoplifting incident to his expired driver's license photo. He spent roughly thirty hours in custody for a crime he had no connection to.

It was the first publicly confirmed case of a false facial-recognition match leading to a wrongful arrest, and the ACLU's subsequent lawsuit produced, in 2024, a landmark settlement barring the Detroit Police Department from arresting anyone based solely on a facial-recognition result or a photo lineup that directly follows one. That settlement is itself the clearest evidence available that the failure was procedural, not computational: officers used a probabilistic investigative lead as if it were a positive identification. Nothing about the technology required that misuse. The department's policies and training permitted it, and so the department bore the liability.

Michigan's MiDAS: automation as a force multiplier for institutional carelessness

Michigan's Integrated Data Automated System, deployed to detect unemployment-insurance fraud, wrongly accused an estimated 40,000 people of fraud between 2013 and 2015, with reported error rates in its fraud determinations running as high as 93 percent. The system flagged trivial data discrepancies as fraud and demanded claimants respond within ten days or face quadruple penalties, wage garnishment, and seized tax refunds — with essentially no human review before punitive action.

Michigan eventually settled the resulting Bauserman v. Unemployment Insurance Agency litigation for $20 million in 2022, covering roughly 3,000 of the wrongfully accused. MiDAS is the purest distillation of the accountability problem: a state agency automated a punitive process, removed the human checkpoint that would have caught obviously flawed fraud flags, and let the system run unaudited for two years while it garnished the wages of tens of thousands of people who had done nothing wrong.

The technology was not sophisticated. The failure was that no one in the chain of command was made to answer for what it did until a class-action lawsuit forced the question seven years later.

Where responsibility actually belongs

Every case above collapses under scrutiny into the same shape: an institution chose to deploy a system, chose not to audit it adequately for the population it would affect, and chose to treat its output as authoritative enough to justify punitive or life-altering action — arrest, deportation-adjacent debt collection, benefit clawback, denial of medical care, rejection from a job.

In every one of these cases, meaningfully better outcomes were available at the time of deployment: Amazon's own engineers caught their bias internally; Obermeyer's fix cost little and worked; Detroit's own post-hoc policy shows facial-recognition matches were always usable as leads rather than proof. The harm was not inevitable. It was the product of institutions choosing speed, cost savings, or the appearance of objectivity over verification.

The accountability framework that follows from this evidence is not complicated, even if implementing it is politically hard. First, liability should attach to the deploying institution, not just the vendor — the hospital that bought Optum's tool, the Belastingdienst, the Detroit Police Department, not merely the engineers who wrote the code, and certainly not the software itself.

Second, vendors should not be permitted to invoke trade-secret protection to block independent audits of systems used in courts, benefits determinations, hiring, or policing; Northpointe's ability to withhold its scoring methodology from the very defendants it was scoring is the single clearest structural failure in the COMPAS story, and it is a policy choice, not a technical necessity.

Third, "human in the loop" cannot satisfy accountability if that human is systematically expected to defer to the system — automation bias, visible in both MiDAS and the Williams arrest, means a rubber-stamping human reviewer adds legal cover without adding actual safety. Fourth, outcome audits disaggregated by race, gender, and other protected characteristics should be a precondition for deployment in any high-stakes domain, not a response undertaken only after journalists or academic researchers force the issue, as happened with ProPublica, Buolamwini and Gebru, and Obermeyer's team.

Europe has legislated a version of the fourth requirement and then deferred it. The EU AI Act's Article 50 transparency duties applied from August 2, 2026, but the obligations on standalone high-risk systems in Annex III — the category covering hiring, benefits, credit scoring and law enforcement — were pushed back to December 2, 2027, and systems embedded in regulated products to August 2, 2028, after the omnibus agreement replaced a conditional trigger with fixed dates. Risk management, documentation and human-oversight duties survived intact; only the clock moved.

None of these four requirements are novel; they are close to the existing malpractice, product-liability, and civil-rights frameworks already applied to other high-stakes decision systems. The only thing standing between the status quo and that framework is a persistent willingness to let "the algorithm decided" function as a legal and moral shield. It should not. The people who built, bought, and deployed these systems made choices. They should answer for them.

Explore

More articles