Skip to content
AI.info

Responsible AI

Risk Appetite, Acceptance Criteria, and Stop Rules

Translate organizational risk tolerance into measurable acceptance criteria, escalation thresholds, and authority to pause or stop an AI system.

By the end you can

Example

22,427 adjudications reviewed, over 93% not fraud

Michigan's Unemployment Insurance Agency ran fraud adjudication through MiDAS with essentially no human review between October 2013 and August 2015. The case for the system was aggregate: determinations at machine speed, and money recovered from them. The harm landed on people with no financial buffer. It landed before anyone checked whether the determination was right. Collection ran through wage garnishment and tax-refund seizure while the check was still outstanding.

The numbers arrived from outside, years later. A federal court recorded in 2018 that “the Michigan Auditor General eventually determined that of the 22,427 robo-adjudications reviewed, over 93% did not involve fraud at all”. The Michigan Supreme Court added the human count in 2022. Roughly 40,000 people had been wrongfully accused of unemployment fraud in that period. An Agency study concluded that approximately 93% of the automated system's fraud determinations were incorrect.

None of that is a modelling failure better statistics would have caught in time. The system ran for the best part of two years because no limit had been written down. Nobody outside the programme could halt it.

  • Benefit metric: Fraud adjudication at machine speed, with recovery pursued on whatever determinations the system produced.
  • Harm metric: The burden fell on people with no buffer, and it was collected before review. Wage garnishment and tax-refund seizure ran ahead of any check.
  • Missing tolerance: No limit was set on how wrong the automated determinations were allowed to be. The Auditor General's figure — over 93% not involving fraud — arrived years afterwards, from outside.
  • Authority gap: Essentially no human review sat between the model and the accusation, for the whole period from October 2013 to August 2015.
  • Decision failure: The two numbers that mattered — 22,427 reviewed, roughly 40,000 people accused — were assembled by a court and an auditor. Not by the operator, and not while the system was running.

What separates accepted risk from unmeasured risk

Risk appetite is the kinds and levels of risk an organization is willing to pursue or retain in service of legitimate objectives. Acceptance criteria turn that position into measurable release conditions. Stop rules keep the authority to pause when the evidence deteriorates. A defensible rule combines quantitative thresholds, qualitative prohibitions, confidence requirements, compensating controls and escalation. It also separates residual risk that leadership knowingly accepted from risk that stayed invisible because nobody measured it.

That distinction is not only an ethical nicety. At least one regulator has made it enforceable, and in the direction organizations least expect. An employer who never collected the required impact data does not get a clean result under the Uniform Guidelines on Employee Selection Procedures. It gets an inference drawn against it. The text is at 29 CFR § 1607.4(D): “Where the user has not maintained data on adverse impact as required by the documentation section of applicable guidelines, the Federal enforcement agencies may draw an inference of adverse impact of the selection process from the failure of the user to maintain such data, if the user has an underutilization of a group in the job category, as compared to the group's representation in the relevant labor market or, in the case of jobs filled from within, the applicable work force.”

Not measuring is not clearance. Under that paragraph the absence of the data is itself evidence.

Silence about a risk reads afterwards as acceptance. The two are not the same: one was chosen against evidence, the other was only never measured. 29 CFR § 1607.4(D) shows a regulator prepared to read the silence against the organization.

Comparison

A stopping boundary fixed before the data existed

An aspirational principle states a value. An acceptance criterion states the evidence a decision needs. A stop rule states something that has to happen tonight.

The third column did its work in the Women's Health Initiative. Its estrogen-plus-progestin arm was a randomised primary-prevention trial of 16,608 postmenopausal women aged 50-79 with an intact uterus. Recruitment ran through 40 US clinical centres between 1993 and 1998. Planned duration: 8.5 years. It did not run that long.

The trial's 2002 report in JAMA says why: “On May 31, 2002, after a mean of 5.2 years of follow-up, the data and safety monitoring board recommended stopping the trial of estrogen plus progestin vs placebo because the test statistic for invasive breast cancer exceeded the stopping boundary for this adverse effect and the global index statistic supported risks exceeding benefits.” The hazard ratios, with nominal 95% confidence intervals, were 1.26 (1.00-1.59) for invasive breast cancer on 290 cases, and 1.29 (1.02-1.63) for CHD on 286 cases. NHLBI announced the halt on 9 July 2002.

Nothing in that was improvised. The boundary was fixed before the data existed. The telemetry was an accumulating test statistic. And the board that recommended the halt did not have to go through the people invested in the trial continuing. That is the whole column, working.

FigureComparison · 3 columns

Aspirational principle

States a desired value.

  • Useful for orientation
  • Usually not measurable alone
  • Does not resolve trade-offs
  • Example: avoid unfair harm

Acceptance criterion

Defines evidence required for a decision.

  • Links metric or scenario to a threshold
  • Names population and uncertainty
  • Can require compensating controls
  • Example: maximum hold disparity and appeal time

Stop rule

Defines when operation must change immediately.

  • Protects against deadline pressure
  • Requires telemetry and authority
  • Can be triggered by incidents or uncertainty
  • Example: pause after unexplained subgroup spike

Visual

The rows of this ladder have names in AI RMF 1.0

Objective, appetite, tolerance, acceptance evidence, stop rule: each row narrows the one above it until a launch becomes decidable. The rows are not an invention of this lesson. NIST named them in the Artificial Intelligence Risk Management Framework, AI RMF 1.0, published in January 2023. Each row there carries an outcome an organization can be held to.

The framework will not fill the tolerance row for you: “While the AI RMF can be used to prioritize risk, it does not prescribe risk tolerance”. What it requires is that the row be filled, and written down. MAP 1.5 is “Organizational risk tolerances are determined and documented”. MEASURE 2.6 asks whether a system's “residual negative risk does not exceed the risk tolerance”. That is the acceptance-evidence row stated as a test. MANAGE 2.4 supplies the authority row, in language stronger than most internal policies manage: “Mechanisms are in place and applied, and responsibilities are assigned and understood, to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use.”

The bottom of the ladder is not a discussion either. Three conditions make a risk level unacceptable: significant negative impacts imminent, severe harms actually occurring, catastrophic risks present. Where an AI system presents them, the framework says “development and deployment should cease in a safe manner until risks can be sufficiently managed.” Cease. Not review.

FigureProcess · 5 steps
  1. 1

    Objective

    The legitimate benefit the system is expected to create.

  2. 2

    Risk appetite

    The categories of risk the organization may accept.

  3. 3

    Tolerance

    The permitted range for a specific metric or scenario.

  4. 4

    Acceptance evidence

    The tests, reviews, and documentation required before release.

  5. 5

    Stop rule

    The trigger and authority for pause, rollback, or retirement.

Steps

Who may pause without commercial permission

Classifying consequences is the cheap step. Naming who may invoke a stop rule without commercial permission is the step that gets negotiated away.

The five steps are not new material. Choosing indicators and setting evidence quality are what MEASURE 2.6 tests when it asks whether residual negative risk exceeds the tolerance. Precommitting decisions is what the Women's Health Initiative did. It wrote its stopping boundary before enrolment, not after the breast-cancer statistic started moving. And the last step is the one MANAGE 2.4 puts in the same breath as the mechanism itself: responsibilities are assigned and understood.

Michigan's Unemployment Insurance Agency had a system that could accuse. From October 2013 to August 2015 it had nobody positioned to stop it.

FigureProcess · 5 steps
  1. 1. Classify consequences

    Separate reversible inconvenience from severe or rights-affecting harm.

  2. 2. Choose indicators

    Measure benefit, error, distribution, workload, recourse, and unknowns.

  3. 3. Set evidence quality

    Require sample coverage, confidence, independent review, and stress tests.

  4. 4. Precommit decisions

    Write launch, restrict, pause, rollback, and retire conditions.

  5. 5. Assign authority

    Name who can invoke the rule without commercial permission.

Example

Forty-five minutes, and no procedure to halt SMARS

Only the last drill produces a fact rather than a document: invoke the emergency stop and time how long it takes. If that sounds like a formality, the SEC has published what the absence of it costs.

On 1 August 2012, while processing 212 small retail orders, Knight Capital's SMARS router sent millions of orders over a 45-minute period. It obtained over 4 million executions in 154 stocks, for more than 397 million shares, and lost over $460 million. The SEC's order the following year names the control that was missing: “Knight also did not have procedures in place to halt SMARS’s operations in response to its own aberrant activity.” Knight had no supervisory procedures for incident response either. The civil money penalty was $12,000,000.

Why that control is the one that goes missing had been set out the year before. Risk checks cost time: “applying risk controls before the start of a trade can slow down an order”, Carol Clark wrote in a Chicago Fed Letter in 2012. Speed is what the business bought. The halt is what speed is traded against. Clark named the missing piece as well — “A “kill switch” that could stop trading at one or more levels”.

  • Tolerance table: Pair each major risk with a threshold, a rationale and an owner. MAP 1.5 asks for tolerances determined and documented, not held in the heads of the people running the system.
  • Uncertainty clause: Specify what happens when the evidence is too weak to estimate the risk. Under 29 CFR § 1607.4(D), missing data counts as evidence against the user rather than as a clean result.
  • Dissent path: Record how a reviewer can escalate disagreement without retaliation. The Women's Health Initiative monitoring board did not have to persuade the investigators before recommending a stop.
  • Emergency action: Invoke the stop and time it. Knight's order records 45 minutes, more than 397 million shares and over $460 million, on a day when no procedure existed to halt SMARS at all.

Key idea

The four-fifths rule disclaims its own number

A threshold is not responsible merely because it is numeric. The value has to reflect consequence, uncertainty, baseline, affected stakeholders, and the organization's capacity to detect and remedy harm.

The best-known numeric acceptance threshold in employment testing says exactly that about itself, in its own text. A selection rate below four-fifths (80%) of the highest group's rate is generally regarded by the federal enforcement agencies as evidence of adverse impact. Four bodies adopted the four-fifths rule jointly in 1978: the EEOC, the Civil Service Commission, the Department of Labor and the Department of Justice. It sits at 29 CFR § 1607.4(D), the same paragraph as the inference above. Then the paragraph disclaims its own number, in both directions. Downward: “Smaller differences in selection rate may nevertheless constitute adverse impact, where they are significant in both statistical and practical terms or where a user's actions have discouraged applicants disproportionately on grounds of race, sex, or ethnic group.” Upward: greater differences may not constitute adverse impact where they are based on small numbers and are not statistically significant.

No acceptance process can prove zero risk. Leadership must record residual risk, dissent, uncertainty, and the reason a benefit justifies continued exposure. A threshold that arrives without that record is a number waiting to be argued about under deadline. A rule written in 1978 was careful not to leave one behind it.

Before approving a threshold, ask what would show it had been set too high, and who is permitted to say so. The drafters of the four-fifths rule answered the first question inside the rule itself.

Case

Article 9(5) asks for acceptable, not for zero

The EU AI Act takes the same position on residual risk. Article 9 says so in three linked paragraphs.

Article 9(2) fixes the cadence. The risk management system is “a continuous iterative process planned and run throughout the entire lifecycle of a high-risk AI system, requiring regular systematic review and updating”. Article 9(8) fixes what an acceptance criterion has to look like. Testing is “carried out against prior defined metrics and probabilistic thresholds that are appropriate to the intended purpose of the high-risk AI system”. Prior defined. The threshold exists before the result does, as the Women's Health Initiative boundary did.

And Article 9(5) does not ask for zero: “The risk management measures referred to in paragraph 2, point (d), shall be such that the relevant residual risk associated with each hazard, as well as the overall residual risk of the high-risk AI systems is judged to be acceptable.” Judged. By someone. That someone should be named in the record. A judgement with no author is the part of an acceptance decision that nobody can later be held to.

The operational conclusion — risk appetite and release authority

An appetite written after launch describes what already happened. It does not constrain it. The Michigan figures — 22,427 robo-adjudications reviewed, over 93% not involving fraud, roughly 40,000 people wrongfully accused — were produced by an Auditor General, a federal court and a state supreme court. All three were working after the fact. Knight's 45 minutes were reconstructed by the SEC a year later. In both records the arithmetic was eventually excellent, and far too late to be a control.

The alternative is on the table in the same lesson. A boundary set before enrolment, and a board authorised to act on it, ended a trial at a mean of 5.2 years against a planned 8.5. AI RMF 1.0 asks under MAP 1.5 that tolerances be determined and documented. Under MANAGE 2.4 it asks that responsibilities to supersede, disengage or deactivate be assigned and understood. Article 9(8) asks for prior defined metrics and probabilistic thresholds. Article 9(5) asks that the residual risk left over be judged acceptable by someone the record can name.

So define what would make risk appetite and release authority require the board to redesign, restrict, remedy or retire this system. Define it in the tense those instruments use. The future.

Key takeaways