Skip to content
AI.info

Research

Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

Overview Research area: AI safety and ethics, specifically the application of systems safety science and sociotechnical analysis to the governance of AI systems (arXiv:2607.14353v1, cs.CY, published 1

arXiv
2607.14353
Published
2026-07-15
Authors
Joshua A. Kroll, Andrew Smart, R. Stuart Geiger, Abigail Z. Jacobs

AI summary

Overview

Research area: AI safety and ethics, specifically the application of systems safety science and sociotechnical analysis to the governance of AI systems (arXiv:2607.14353v1, cs.CY, published 15 Jul 2026).

Technical level: Intermediate. The paper contains no mathematics, code, or experimental results, so it is readable by non-specialists, but it presumes familiarity with current AI safety debates (alignment, benchmarks, red teaming, auditing) and with the safety science literature it draws on.

Scope: A conceptual position paper arguing that AI safety must be treated as a system-level sociotechnical problem rather than a set of component-level technical fixes.

What This Paper Is About

The paper argues that the AI field is repeating mistakes that safety science has already diagnosed in disasters such as Chernobyl, Three Mile Island, Fukushima-Daiichi, Bhopal, and the Challenger explosion. Those catastrophes were widely mistaken for unforeseeable freak accidents, but closer examination shows the hazards were known beforehand and were not acted upon because of social, structural, political, and economic factors. The authors' goal is to translate those "unlearned lessons" about organizational culture, incentives, risk perception, and responsibility into concrete guidance for how AI systems are designed, evaluated, and governed.

Key Contributions

  1. A framework for applying systems safety to AI. The paper defines an "AI system" as a deployed, end-to-end sociotechnical system comprising technical subsystems (data collection and quality controls, data pipelines, feature generation, model training workflows, deployment infrastructure, user interfaces, logging and auditing, benchmarks and evaluation suites) and human/organizational processes (problem formulation, requirements analysis, metric selection and objective setting, review and approval, stakeholder engagement, procurement, individual and institutional incentives, cultural practices, and political-economic forces).

  2. A high-level taxonomy of "unlearned lessons." Six categories of organizational failure etiology are derived from disaster literature and paired, in Table 1, with symptoms, historical examples, and their expression in modern AI systems: Poor Risk Perception; Incentives; Permanent Rush Cultures; No Bad News; Safe Components Do Not Imply Safe Systems; and Kick the Can.

  3. A critical reframing of mainstream AI safety approaches. The authors characterize existing AI safety research as aiming either to increase component reliability (better models, more benchmarks, bias metrics), to enhance component performance requirements (alignment, fine-tuning, robustness), or to perform what they call "unvalidated risk-reduction rituals" such as model-focused auditing and documentation. They argue these treat safety as a "solution" rather than as continual effort, management, and repair.

  4. A sociotechnical account of mitigation. Section 4 begins outlining how safety cultures, curmudgeons and critics, traceable processes, psychological safety, diverse teams, meaningful external participation, and treating failure as normal and expected can improve—but never solve—AI safety. The provided text is truncated partway through Section 4.3, so the mitigation framework is only partially visible.

Main Findings

  • Disasters were not freak accidents. The authors state that risks and hazards in events such as Chernobyl, Three Mile Island, Fukushima-Daiichi, Bhopal, and Challenger were well-known beforehand but were not acted upon due to social structural, political, and economic factors.

  • Safety is an emergent property. System behavior derives from the structure and hierarchy of components, their interactions, and the communication and control flows between them. Accidents can occur even when no component has individually failed and when all components follow their functional specifications.

  • Component verification and validation is necessary but insufficient. Verification ("did we build the system right") and validation ("did we build the right system") evaluate behavior against specification under assumed conditions; neither captures how a composed system behaves in real-world use.

  • Challenger as a canonical case. Diane Vaughan's investigation identified overlapping factors producing a "normalization of deviance," and the paper notes that knowledge of the O-ring failure mode was distributed across NASA and contractor teams but obscured by social, cultural, organizational, political, and economic conditions.

  • High benchmark performance obscures system-level risk. The paper uses the AUC (Area Under the Curve) benchmark metric as a symbol of this problem, arguing that high performance does not establish safety.

  • The healthcare allocation example. A widely-cited healthcare machine learning model allocated care for high-risk patients using actual healthcare costs as a proxy for medical severity. Because Black patients spend far less on healthcare for equivalent conditions due to structural inequalities in the US, the model systematically underestimated their health needs—even though the developers were explicitly concerned about bias, they failed to perceive that their choice of proxy obscured the risk.

  • Incentives displace safety work. Business pressure caused Boeing to suppress safety activities around a new subsystem in the 737 MAX-8, arguing to regulators that performance in the existing Boeing KC-46 air tanker was sufficient evidence, while key differences relevant to the MAX-8 were not communicated to pilots; this led to fatal accidents and grounding of the fleet for well over a year.

  • Permanent rush cultures sacrifice safety margins. The USS Scorpion was the only SUBSAFE-covered ship lost after its planned maintenance overhaul was reduced in scope and time, eliminating safety-critical work to return it to service faster under Cold War pressures. In the former Soviet Union, Politburo demands to accelerate nuclear plant construction led the industry to ignore hazards and design flaws; near the Chernobyl accident, a planned safety test was delayed many hours into the night to keep power output high for end-of-quarter production quotas at nearby factories.

  • "No bad news" cultures suppress risk information. Silicon Valley Bank's risk department ran an internal stress test revealing vulnerability of its bond portfolio to rising interest rates; the finding led the risk organization to change core assumptions in the bank's risk model relating deposit income to interest rate changes, and approximately 2.5 years later, bond sales at a loss tied to rising rates were the proximate inciting factor in a run that led to the bank's collapse. At Meta, the chief operating officer named her personal conference room "Only Good News," and former employee Frances Haugen testified that internal research on Instagram's engagement algorithms and mental health impacts on teenage girls was systematically ignored.

  • Responsibility is displaced onto operators—the "moral crumple zone." In the 2018 Uber automated driving crash that killed a bicyclist, the safety driver was charged with negligent homicide even though the automated system also failed to identify the bicyclist. The paper also points to Chernobyl and Air France flight 447, where Elish argues the only systemically possible "malfunction" is the human pilot, and cites Dekker and Woods' position that safety improves not by assigning functions to whichever of human or automation is better, but by helping the two approaches to control "get along."

  • Culture is not an operating system. Citing Silbey, the authors warn that engineers often treat culture as an instrumental technology to be installed. Culture is a historically produced, power-laden field of norms and meanings that conditions which safety actions are intelligible and legitimate—it is a permitting frame, not a mechanism.

  • Human backups to automation are generally ineffective. Citing Bainbridge, the paper notes that not being in active control leads to complacency and degradation of situation awareness.

Methodology in Plain English

This is a conceptual synthesis, not an empirical study. The authors review primary and secondary literature from the systems safety tradition and from documented sociotechnical disasters across aviation, nuclear power, industrial manufacturing, finance, and medicine. From that literature they derive a high-level taxonomy of recurring organizational failure modes—the "unlearned lessons" of the title. They then map each category onto contemporary AI research and development practice, illustrating how the same failure mode appears in benchmark-driven evaluation, release incentives, "AI race" framing, and human-in-the-loop liability shifting. A companion table aligns each lesson with its symptoms, its historical disaster examples, and its AI-era expression. They deliberately avoid specifying a particular AI system architecture or lifecycle model, stating that system-level risk reasoning should not depend on such details. The mitigation section organizes recommendations around the safety culture literature, including research on high-reliability organizations. No quantitative experiments, datasets, or benchmark evaluations are reported.

Why This Matters

Impact on research: The paper challenges the dominant framing in AI safety, in which better models, more benchmarks, alignment techniques, guardrails, and interpretability are treated as the path to safety. It argues that these component-level, extensionally visible properties are necessary for responsible design but do not on their own establish safety, and that treating them as sufficient launders responsibility onto developers or end users. It calls for safety to be treated as a first-order engineering concern that includes social and organizational dynamics.

Real-world applications:

  • Healthcare: The care-allocation model that used healthcare costs as a proxy for medical severity shows how an unintended proxy variable can encode structural inequality despite developers' explicit concern about bias.
  • Automated driving: The 2018 Uber crash illustrates how automated perception failures plus human "safety driver" fallback arrangements produce moral crumple zones rather than safety.
  • Aviation: The Boeing 737 MAX-8 case shows how commercial pressure and reliance on evidence from a different operational context can suppress safety analysis.
  • Finance: The Silicon Valley Bank collapse shows how changing a risk model's core assumptions rather than acting on its findings can convert a known vulnerability into a catastrophe.
  • Infrastructure and software: The paper notes that improperly formatted updates to security scanning data in a widely used security scanning tool caused a large fraction of the world's computers to fail-stop and require rebooting, disrupting transportation networks, financial markets, and other critical infrastructure.

Industry relevance: The paper targets organizational practice directly. It recommends dedicated high-level roles for safety personnel, cultural acceptance of and organizational incentives for curmudgeons (pathologically critical thinkers), critical self-evaluation after safety violations or near misses, traceable processes, and more stakeholder-centered orientations and incentives. It also observes that the assumption of machine learning as an entirely novel form of software engineering has led to widespread abandonment of traditional requirements engineering.

Future Directions

  • How to build safety cultures without instrumentalizing them. The authors explicitly warn that a safety culture is not a deployable artifact with predictable risk-reduction effects, leaving open how organizations can cultivate one nonetheless.
  • How to make safety concern decisive rather than advisory. The paper notes that being permitted to raise a concern is not the same as avoiding the outcome; the open question is whether an organization has agreed beforehand to treat the concern as decisive rather than relitigating it each time.
  • How to trace requirements and responsibilities through AI supply chains. The abstract names traceability of requirements and responsibilities as one of the three areas where AI can benefit from unlearned lessons, and the "Kick the Can" analysis raises unresolved questions about where liability should vest when human-in-the-loop arrangements shift risk to users and contractors.
  • How to improve organizational risk perception, communication, and analysis. The abstract identifies this as a core area for improvement, and the taxonomy shows it cuts across all six unlearned lessons.

Target Audience

AI researchers and engineers who design, evaluate, or deploy learned models; AI safety and AI ethics researchers looking for a systems-level alternative to component-focused framings; engineering and product leaders responsible for organizational incentives, review processes, and release cadence; regulators, policymakers, and procurement officials who must judge safety claims; and scholars in safety science, science and technology studies, HCI, and CSCW who study how technical work is embedded in organizational and political contexts. The paper is written by authors affiliated with the Naval Postgraduate School, Google Research, the University of California San Diego, and the University of Michigan, Ann Arbor.

Authors’ abstract

As automated decision-making and data-driven technologies pervade society and are used to manage consequential outcomes, understanding the technology's capabilities, limitations, and attendant risks in context requires analysis of full sociotechnical systems. Sociotechnical analysis of risks in highly complex systems provides clear lessons for the design and evaluation of AI systems, transcending a technical focus on reliable or "responsibly designed" components to understand risks at a systems level. Human-made catastrophes have been studied for decades because of the severity of these events: consider Chernobyl, Three Mile Island, Fukushima-Daiichi, Bhopal, the Challenger disaster. A common misconception is that these kinds of events are freak accidents, resulting from the inherently unforeseeable interactions in complex systems. Closer examination reveals that the risks and hazards were well-known beforehand but not acted upon due to social structural, political and economic factors. We outline several areas where the development and use of AI can benefit from learning these unlearned lessons: improved risk perception, communication, and analysis at the organizational level; traceability of requirements and responsibilities; and holistic approaches to responsibility and safety that include social and organizational dynamics as first-order engineering concerns. For each area, we offer concrete unlearned lessons and exemplify how they led to failure in prior accidents as well as examples of how these lessons remain unlearned for modern computing systems, particularly AI.

Read the original paper