Skip to content
AI.info

Responsible AI

Policies, Standards, Controls, and Evidence

Distinguish principles, policies, standards, procedures, controls, evidence, and exceptions in a responsible AI management system.

By the end you can

Key idea

A record is not proof the control ran — and audit standards say so

A record is not evidence that the control it describes actually operated. Screenshots, checklists and signatures turn into compliance theatre the moment nobody can tie them to a system version, the test data, the reviewers, the exceptions, and what happened next. A record proves a control ran only when it carries a version, a reviewer and a date.

Audit practice drew that line long before AI, and wrote it down. Design and operation are two different things, and they get tested separately. The SEC said so on 27 July 2007, approving Auditing Standard No. 5 in place of the standard before it: the auditor must “evaluate and test both the design and the operating effectiveness of internal control”. That standard is now PCAOB AS 2201, and it still keeps the two apart — design effectiveness at ¶.42, operating effectiveness at ¶.44.

A document cannot discharge the second. Paragraph .44 asks something no document answers: “The auditor should test the operating effectiveness of a control by determining whether the control is operating as designed and whether the person performing the control possesses the necessary authority and competence to perform the control effectively.” Read the last clause again. It asks about a person.

Then ¶.46 refuses to let the amount of evidence follow from the control's existence: “the evidence necessary to persuade the auditor that the control is effective depends upon the risk associated with the control”. The riskier the control, the more proof it owes. Existing is not proof.

Controls conflict, decay, and create new burdens. Whoever runs the management system has to test whether they still work, and retire the ones that no longer manage the risk they were built for.

Counting artifacts cannot catch a control that quietly decayed; AS 2201 ¶.44 asks a harder question — did the control operate as designed, and did the person performing it have the authority and competence to perform it.

Example

One untested system: Rite Aid's facial recognition

A commitment that is never converted into a tested control leaves nothing behind but the commitment.

On 19 December 2023 the Federal Trade Commission sued Rite Aid over its use of facial recognition in stores. Paragraph 5 of the complaint lists the absent controls one after another: “Among other things, Rite Aid failed to consider or address foreseeable harms to consumers flowing from its use of facial recognition technology, failed to test or assess the technology’s accuracy before or after deployment, failed to enforce image quality standards that were necessary for the technology to function accurately, and failed to take reasonable steps to train and oversee the employees charged with operating the technology in Rite Aid stores.”

Four alleged failures. Not one of them is a failure of the model.

The stipulated order for permanent injunction was entered on 26 February 2024. Provision I bans any Facial Recognition or Analysis System for five years. Provision III writes the missing control card into an enforceable obligation: a documented monitoring program, a designated responsible employee, and a written System Assessment at least once every twelve months. Read the ladder below against that case. The regulator had to supply, by order, the rungs the company never built.

  • Principle: A broad value is stated — here, that foreseeable harms to consumers from a deployed technology will be considered and addressed. Paragraph 5 of the FTC's complaint alleges the failure begins at exactly that rung.
  • Policy: Leadership declares mandatory organizational expectations. A declaration on its own leaves an investigator nothing to inspect. What the complaint examines is not what was announced but what was performed.
  • Standard: A defined method or threshold constrains implementation. Rite Aid, the FTC alleged, “failed to enforce image quality standards that were necessary for the technology to function accurately” — the threshold the whole system's accuracy rested on.
  • Control: A mechanism prevents, detects, or corrects a specific failure. Rite Aid “failed to test or assess the technology’s accuracy before or after deployment”, the FTC alleged, and failed to train and oversee the employees operating it. No prevention, no detection, nobody positioned to correct.
  • Evidence: A record shows whether the control operated for the relevant release. Provision III of the 26 February 2024 order supplies the format the company lacked: a documented monitoring program, a designated responsible employee, a written System Assessment at least once every twelve months.

Visual

Down the ladder from value to testable record

Principle, policy, standard, procedure, control: each rung down the ladder turns something intended into something a reviewer can test. Only the bottom rung can be put on trial. And even there the trial splits in two. AS 2201 asks first whether the control was designed to do the job (¶.42). Then it asks whether the control actually ran, in the hands of someone with the authority and competence to run it (¶.44).

FigureProcess · 5 steps
  1. 1

    Principle

    Value or aspiration that guides judgment.

  2. 2

    Policy

    Mandatory organizational direction and accountability.

  3. 3

    Standard

    Consistent rule, method, or minimum requirement.

  4. 4

    Procedure

    Repeatable steps used by a role or team.

  5. 5

    Control and evidence

    Mechanism that manages risk and the record that demonstrates operation.

Ten fields that turn a value into a check

Governance works down a ladder. At the top is what an organization intends; at the bottom is what somebody actually does. Principles express values. Policies set mandatory direction. Standards define consistent requirements. Procedures describe work. Controls manage specific risks. Evidence shows whether the control operated.

A control gets ten fields. An objective, an owner, a trigger, a frequency. An input, a method, an expected result. Retained evidence, an exception path, a test procedure. The evidence has to correspond to the deployed version and context, not to a generic template.

Those fields are not a local invention. The GAO published the same idea as a catalogue in 2021, pairing every numbered key practice with questions to consider, audit procedures, and the types of evidence an auditor should collect.

The owner field is not administrative decoration. AS 2201 ¶.44 makes the person performing the control part of what is tested: does that person have the necessary authority and competence to perform it effectively. A blank there is not a formatting defect. It is the point at which the check stops binding anyone.

Ten fields sounds heavy until three product teams read the same fairness policy and none of them can name who owns the check.

Only the named owner turns a policy value into something a person can be held to — and AS 2201 ¶.44 tests whether that person has “the necessary authority and competence to perform the control effectively”.

Comparison

Preventive control, Detective control, or Corrective control?

Preventive controls reduce the chance of a failure. Detective controls find it. Corrective controls decide how bad it gets. The split is not a classroom scheme. It is standing public doctrine in two independent standards.

Timing is what separates the first pair. The GAO's Green Book, its standards for internal control in the federal government, sets it out: “Control activities can be either preventive or detective. The main difference between preventive and detective control activities is timing, that is, when the control activity occurs within an entity’s operations. A preventive control activity is designed to avoid an unintended event or result before it occurs. A detective control activity is designed to discover and timely correct an unintended event or result after it occurs.”

That definition has survived a rewrite. It stood at ¶10.04 of the 2014 edition. It stands at ¶10.10 of the revision published on 15 May 2025, effective from fiscal year 2026. Both editions ask for the two to be designed. Not one.

Financial reporting has the same pair, at AS 2201 ¶.A8: “Controls over financial reporting may be preventive controls or detective controls. Effective internal control over financial reporting often includes a combination of preventive and detective controls.” That second sentence is the portfolio argument in one line. Not a type. A combination.

The corrective work of the third column is already lodged inside the Green Book's detective definition, which asks the activity not merely to discover but to timely correct.

FigureComparison · 3 columns

Preventive control

Reduces the chance that a failure occurs.

  • Examples: access restriction, prohibited-use gate
  • Acts before the harmful event
  • May create operational friction
  • Needs bypass monitoring

Detective control

Finds a failure or deteriorating condition.

  • Examples: drift alert, complaint analysis
  • Depends on observable signals
  • Can suffer delay and false alarms
  • Needs response ownership

Corrective control

Limits harm and restores a safe state.

  • Examples: rollback, correction, compensation
  • Acts after a failure or incident
  • Requires recovery capacity
  • Should feed lessons into prevention

Example

Control cards, traceability, and control debt

Three of the four drills below already have a published format you can copy. The fourth produces the uncomfortable total: how many controls exist on paper and cannot currently be tested at all.

The format is GAO-21-519SP, the accountability framework for AI published on 30 June 2021. It sorts four principles — governance, data, performance, monitoring — into 31 numbered key practices, running from 1.1 to 4.5. Its first page says what each practice carries: “Each practice includes a set of questions for entities, auditors, and third-party assessors to consider, along with audit procedures and types of evidence for auditors and third-party assessors to collect.” A practice got into that catalogue only when at least two independent sources called it important.

  • Control card: Document objective, owner, trigger, procedure, evidence, and failure response — the shape GAO already prints for each of its 31 key practices. Practice 1.2, on roles and responsibilities, carries all four parts: practice, questions to consider, audit procedures, types of evidence.
  • Traceability test: Link one policy statement to a standard, a control, an artifact, and a monitored outcome. GAO does not leave this to good intentions either. Traceability is its own numbered key practice, 4.3, with its own audit procedures.
  • Exception review: Identify who may approve a deviation and what compensating control is required. AS 2201 ¶.46 sets the scale of what the answer must carry: “the evidence necessary to persuade the auditor that the control is effective depends upon the risk associated with the control”.
  • Control debt: List the controls that exist on paper but cannot currently be tested. Rite Aid's facial-recognition deployment sat in exactly that state: no accuracy testing before or after deployment, no training for the staff acting on it. It stayed there until the order of 26 February 2024 imposed a designated responsible employee and a written System Assessment at least once every twelve months.

Steps

Design the control from the failure, not the policy

Every control here begins from a failure mechanism. A control designed from a policy sentence tends to protect the sentence.

The five steps below are not a private scheme. They are the order in which two published catalogues already work. The GAO framework moves from a numbered key practice, to the questions an entity, auditor or third-party assessor should consider, to audit procedures, to the types of evidence to collect. Objective, then operation, then record. AS 2201 tests the result in two passes that must not be collapsed: design effectiveness at ¶.42, operating effectiveness at ¶.44. And at ¶.46 the amount of evidence demanded is set by the risk attached to the control, not by the fact that the control exists.

Step 5 is where most control sets fail. It is the only step that can return a verdict the organisation did not want.

FigureProcess · 5 steps
  1. 1. Start from the risk

    Write the failure mechanism and consequence the control should address.

  2. 2. Define the objective

    State the condition the control must create or preserve.

  3. 3. Design operation

    Assign owner, timing, inputs, method, exception, and escalation.

  4. 4. Specify evidence

    Record what proves operation for a particular version and context.

  5. 5. Test effectiveness

    Use samples, re-performance, incidents, and outcome data to challenge the control.

A control with no evidence is an intention

A control with no evidence that it ever operated is an intention. Auditors have a shorter word for it.

They also have a precise term for the other case, the control that was designed properly and quietly stopped running. The Green Book's 2025 revision fixes it at OV3.06: “A deficiency in operation exists when a properly designed control does not operate as designed”. That one sentence is why design and operation are tested separately. It is also why a well-drafted control card proves nothing on its own.

Decide in advance what a failed control test triggers: redesign, restrict, remedy, or retire. Rite Aid's order shows the outer end of that range. Provision I removed the technology for five years. Provision III specified what would have to exist before anything comparable could operate again.

Case

What “certifiable” rests on: ISO/IEC 42006:2025

Most control sets are now written against one of two external documents. The first is ISO/IEC 42001:2023, published on 18 December 2023 — the first management-system standard for artificial intelligence.

It is certifiable. And certifiability is not a bare property of the standard. It is a second standard, governing the certifiers. ISO/IEC 42006:2025, published on 7 July 2025, sets requirements for the bodies that audit and certify an AI management system against ISO/IEC 42001, adding to the general certification-body standard ISO/IEC 17021-1. Its abstract states the purpose plainly: “The requirements contained in this document, when implemented, support the demonstration of competence, consistency and reliability by the bodies performing auditing and certification of an artificial intelligence management system (AIMS) according to ISO/IEC 42001 for organizations that provide, develop or use AI systems.”

The certificate is itself a controlled process, with a dated rulebook behind it.

Case

NIST’s framework, eleven months earlier and voluntary

The other document arrived eleven months earlier and asks for nothing in return. NIST released the AI Risk Management Framework 1.0 on 26 January 2023. It carries no certificate at all, and it says so itself, in the executive summary: “The Framework is intended to be voluntary, rights-preserving, non-sector-specific, and use-case agnostic, providing flexibility to organizations of all sizes and in all sectors and throughout society to implement the approaches in the Framework.”

Voluntary does not mean shapeless. The AI RMF Core is a fixed structure: 4 functions — GOVERN, MAP, MEASURE, MANAGE — broken into 19 categories and 72 numbered subcategories. The numbering is precise enough to write a control against. GOVERN 1.5 covers ongoing monitoring and periodic review. GOVERN 1.7 covers decommissioning AI systems.

Choosing between the two documents is a governance decision, not a procurement one. One gives you an auditable certificate with certifier requirements behind it. The other gives you 72 addressable subcategories and no certificate.

Key takeaways