Skip to content
AI.info

Responsible AI

Responsible AI as Sociotechnical Governance

Build a systems-level view of responsible AI that connects models, workflows, institutions, people, and consequences.

By the end you can

Example

The same detection logic, two workflows, 85 percent against 44 percent

Michigan's unemployment agency ran a fraud-detection programme and then reviewed itself. It pulled every fraud-penalty case from October 2013 to August 2015 that was never appealed — 62,784 of them — and sorted them by how the decision had actually been reached.

40,195 had been resolved by computer program alone. 85 percent of those fraud findings were reversed.

22,589 had been initiated by the computer and then referred to a human investigator. 44 percent of those were reversed.

The agency began refunding more than $20.8 million. Its own press release, on 11 August 2017, states the first half of the comparison directly: “Of those cases, 40,195 were originally resolved by way of computer program based on available information. As part of the review, 85 percent of these original fraud findings were reversed.”

No retraining, no new features, no change of threshold separates those two figures. What separates them is whether a person stood between the program and the finding.

  • Model result: The detection logic is the constant across both queues. Whatever its offline quality, it is not the variable that moved between an 85 percent reversal rate and a 44 percent one.
  • Institutional process: The review covers only the fraud-penalty cases from October 2013 to August 2015 that were never appealed — 62,784 of them. The population the numbers describe is the one that did not contest a finding.
  • Decision rule: 40,195 cases were resolved by computer program alone. 22,589 were initiated by the computer and then referred to a human investigator. The routing rule is the governed object here, not the predictor.
  • Affected people: The findings landed on people who sat outside every design and evaluation decision. More than $20.8 million began to be refunded only after the agency reviewed itself.
  • Governance gap: Cahoo v. SAS Institute recites that the system made all fraud determinations between October 2013 and August 2015 with no human review. The opinion opens by describing a system that “lacked human oversight”.

Visual

Five layers between intent and the affected world — and which one moved

Five layers stand between an intended decision and the people who live with its output: purpose and context, technical components, operational workflow, institutional environment, and affected world. Most layer diagrams are an argument. Michigan's review is a measurement.

One detection logic ran through two operational workflows from October 2013 to August 2015. The agency reviewed 62,784 never-appealed cases. Among the 40,195 resolved by computer program alone, 85 percent of the original fraud findings were reversed. Among the 22,589 the computer initiated and then referred to a human investigator, 44 percent were reversed. The technical-components layer accounts for none of the distance between those two numbers.

A layer diagram is a claim about where the variance lives. Here it lived in the workflow layer. It was measured, it was refunded at more than $20.8 million, and it was recited in court as a system that made every fraud determination with no human review.

FigureProcess · 5 steps
  1. 1

    Purpose and context

    The intended decision, population, consequence, and legal or professional setting.

  2. 2

    Technical components

    Data, features, model, thresholds, retrieval, tools, and infrastructure.

  3. 3

    Operational workflow

    People, queues, handoffs, overrides, escalation, and service capacity.

  4. 4

    Institutional environment

    Policies, incentives, budgets, vendors, and accountability structures.

  5. 5

    Affected world

    People, communities, markets, ecosystems, and democratic processes touched by outcomes.

The unit of governance is the decision system

Responsible AI means managing risks, impacts and accountability with discipline across an AI system's full social and technical context. A model can be accurate while the process around it stays unfair, unsafe, unlawful or ineffective. So the unit of governance is the decision system: objectives, data, model, interface; the users and affected people it reaches; and the policies, incentives, monitoring, recourse and retirement that hold it in place.

A court has drawn the unit exactly there. SyRI was the Dutch welfare-fraud risk-profiling system. On 5 February 2020 the District Court of The Hague struck down the legislation that authorised it. The judgment's own summary heading reads “SyRI legislation in breach of European Convention on Human Rights”. The court holds: “For this reason, the court declares in this judgment that Section 65 SUWI Act and Chapter 5a SUWI Decree have no binding effect, being contrary to Article 8 paragraph 2 ECHR.”

No accuracy figure appears anywhere in that finding. The arrangement was insufficiently transparent and verifiable, and that was enough to void it. Rachovitsa and Johann, in the Human Rights Law Review, read the judgment the same way: the case turns on the Article 8(2) breach and the declaration of no binding effect.

Technical controls matter. They cannot make the institutional choices about who receives benefits, burdens, voice and remedy. People make those.

Draw the unit at the model and you can never see what The Hague saw: an entire instrument voided on transparency and verifiability, with not one accuracy figure in the judgment.

Comparison

Model governance, System governance, or Symbolic ethics?

Model governance, system governance and symbolic ethics all travel under the word responsible. They cover very different amounts of the system.

Model governance controls the model artifact and its measured behaviour — version, evaluation, limits and monitoring, the model card and the release test. It is useful. It is also narrower than the deployed system, because it cannot observe harms outside its data boundary. That boundary has been measured. In Science in 2019, Obermeyer and colleagues dissected a commercial health-risk algorithm that reached millions of patients. It was accurate on the label it had been built to predict: health care cost. It produced large racial bias in the thing the programme actually allocated. Their abstract: “At a given risk score, Black patients are considerably sicker than White patients, as evidenced by signs of uncontrolled illnesses. Remedying this disparity would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%.” Every artifact-level test the algorithm passed, it passed honestly. The harm sat in the choice of target variable, and no measurement of the artifact can report that.

System governance controls the complete decision workflow — model, people, policy and operations together. It tracks downstream consequences and recourse, and it gives someone the authority to stop or redesign. It is the level at which 17.7 against 46.5 percent, or Michigan's 85 against 44, is a governable fact rather than a finding published after the damage.

Symbolic ethics states broad values with no enforceable mechanism. It can align language and aspiration. It often lacks owners and evidence, and it can conceal unresolved trade-offs: a principle statement with no release gate behind it.

FigureComparison · 3 columns

Model governance

Controls the model artifact and its measured behavior.

  • Version, evaluation, limits, and monitoring
  • Useful but narrower than the deployed system
  • Cannot observe harms outside the data boundary
  • Example: model card and release test

System governance

Controls the complete decision workflow.

  • Includes model, people, policy, and operations
  • Tracks downstream consequences and recourse
  • Assigns authority to stop or redesign
  • Example: impact assessment and operating controls

Symbolic ethics

States broad values without enforceable mechanisms.

  • Can align language and aspiration
  • Often lacks owners and evidence
  • May conceal unresolved trade-offs
  • Example: principle statement without release gate

Analogy

A navigation system inside a transport network

Judging a city transport network by the accuracy of its route planner is a familiar mistake. The planner calculates fast routes. Fares, station access, schedules and service cuts decide who can actually travel.

A route planner reads the network. A deployed AI system changes it. People reroute around the recommendation, the operator reschedules to match the new demand, and the next model is trained on that rearranged city.

The rearrangement reaches the evidence too. Michigan's review could only examine the fraud-penalty cases that were never appealed — 62,784 of them. Appeal is itself a route through the system. It is available to some people and not to others, and it decides which findings are ever revisited.

A capable component cannot make an inequitable system responsible by itself.

Key idea

What a written principle proved when somebody went and counted

A responsible-AI principle does not become a control merely because it appears in a policy. Someone has to translate it into decision rights, evidence requirements, monitoring, and consequences for non-compliance. No framework can eliminate every conflict among values or predict every social effect. Whoever holds the decision rights must preserve uncertainty, dissent, and the option not to deploy.

New York City's Local Law 144 is the first mandate anywhere requiring bias audits for automated employment decision tools. Then somebody went and counted what it produced. The count came out in 2024, from researchers at Cornell, the Data & Society Research Institute and Consumer Reports: “In this study, 155 student investigators recorded 391 employers' compliance with LL 144 and the user experience for prospective job applicants. Among these employers, 18 posted audit reports and 13 posted transparency notices.” Nearly all of the posted audits reported an impact factor over 0.8.

The duty here was not aspirational. It was statutory, it was specific, and the required evidence had to be posted where anybody could read it. What was missing was the boundary work. Nobody had been assigned to establish who was in scope, so the obligation existed without ever attaching to a named set of employers. That is the gap a document cannot close about itself. The paper's name for it is null compliance.

A binding statute with a public-posting requirement yielded 18 findable audit reports across the 391 employers checked. The principle was law and still not a control.

Case

Why govern runs through the other three

NIST published its Artificial Intelligence Risk Management Framework in January 2023. It states that “the Core is composed of four functions: GOVERN, MAP, MEASURE, and MANAGE”.

Where governance sits is not left to interpretation. A figure caption in the framework says it: “Governance is designed to be a cross-cutting function to inform and be infused throughout the other three functions.” Infused throughout. Not a gate cleared once before the technical work begins, and not a separate jurisdiction with its own owner.

The framework gives the reason: “AI systems are inherently socio-technical in nature, meaning they are influenced by societal dynamics and human behavior.” If risk emerges from the interplay of technical aspects and societal factors, a governance function that stopped at the artifact would be governing the smaller half. So the risk the framework asks about is never the model alone. It is the model together with whoever operates it, the institution around it, and the place it lands. That is another way of stating the difference between 85 percent and 44 percent.

Steps

Name the decision first, write the stop rule last

The sequence opens by naming the decision and closes by writing the conditions under which the system stops: define the decision; draw the system boundary; map impacts and power; assign accountable owners; set evidence and stop rules.

For some deployers the first three steps are no longer a matter of house style. They are law. The EU AI Act was published on 12 July 2024 and applies from 2 August 2026. Before a public body turns a high-risk system on, it has to assess what that system will do to people's fundamental rights. Article 27(1) reads: “Prior to deploying a high-risk AI system referred to in Article 6(2), with the exception of high-risk AI systems intended to be used in the area listed in point 2 of Annex III, deployers that are bodies governed by public law, or are private entities providing public services, and deployers of high-risk AI systems referred to in points 5 (b) and (c) of Annex III, shall perform an assessment of the impact on fundamental rights that the use of such system may produce.”

Six elements are mandated in that assessment. Three of them are steps 3 and 4 of this sequence written as law: the categories of people likely to be affected, the specific risks of harm to them, and the internal governance and complaint mechanisms to be used if those risks materialise. Annex III point 5(a) covers AI used by public authorities to evaluate eligibility for essential public assistance benefits and services. That is the same family of decision as a benefits agency's fraud determination. Naming who is affected and where they can complain is required content, not good practice.

FigureProcess · 5 steps
  1. 1. Define the decision

    Name what changes for whom when the system output is accepted.

  2. 2. Draw the system boundary

    Include data sources, people, vendors, policies, and downstream processes.

  3. 3. Map impacts and power

    Identify benefits, burdens, voice, dependence, and available remedy.

  4. 4. Assign accountable owners

    Separate delivery, risk challenge, approval, monitoring, and audit.

  5. 5. Set evidence and stop rules

    Specify what must be shown before launch and what triggers restriction or retirement.

Example

What the review has to produce

Four drills, each ending in something a reviewer can hold: a boundary sketch, a power map, a list of gaps, and one written condition for not building at all.

The fourth is not this lesson's invention. NIST makes it an output of the MAP function: “After completing the MAP function, Framework users should have sufficient contextual knowledge about AI system impacts to inform an initial go/no-go decision about whether to design, develop, or deploy an AI system.” The framework lists among its own benefits “explicit processes for making go/no-go system commissioning and deployment decisions”. A national standards body writes the option of not building into the process. A team with no such condition on paper has skipped a documented step, not exercised restraint.

  • Boundary sketch: Draw the system from data collection to the final human or automated decision, and mark the routing choices inside it. It was the split between computer-resolved and investigator-referred cases that moved Michigan's reversal rate.
  • Power map: List who can approve, override, appeal, exit, and absorb the cost of failure. Appeal is a boundary of the evidence as well as a right: Michigan's review of 62,784 cases is a review of the cases nobody appealed.
  • Evidence gap: Mark each claim that rests on assumption rather than measured evidence. That includes the claim that a required document exists somewhere — exactly what 155 student investigators had to go out and check for 391 employers.
  • Negative option: Write one condition under which the responsible choice is not to use AI. Write what happens after launch too. MANAGE 4.1 requires that “Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management.”

Governance nobody can see is a document

Governance nobody can see, measure or correct is a document. Michigan ran a fraud-detection programme with no human review from October 2013 to August 2015. It took a review of 62,784 never-appealed cases to establish that 85 percent of the findings made by computer program alone were reversed, and that more than $20.8 million had to go back. New York City had a statute, public postings and a named audit obligation; 155 student investigators found 18 posted audit reports across 391 employers. The Hague had a piece of legislation and voided it on transparency and verifiability. In none of the three did the missing thing turn out to be a better model.

So write it down in advance: when does sociotechnical governance require the team to redesign, restrict, remedy or retire the system? NIST puts a go/no-go decision at the end of MAP. MANAGE 4.1 requires appeal, override and decommissioning after deployment. Article 27 of the EU AI Act makes the affected categories and the complaint mechanisms mandatory content for public-service deployers.

Each of those is somebody's job on a named day. A principle is not.

Key takeaways