AI literacy basics
Capstone: Audit an AI System End to End
Apply the course framework to a support assistant that classifies requests, retrieves policy, drafts replies, and proposes actions under human review.
By the end you can
- Map an AI system across objective, tasks, data, models, controls, people, and outcomes
- Identify technical, operational, and sociotechnical failure modes
- Select evidence, baselines, safeguards, and stopping conditions for a pilot
- Write a bounded proceed, narrow, postpone, or reject recommendation
Example
The HarborHelp proposal
HarborHelp is a proposed assistant for a subscription-software support team. It will classify incoming messages, retrieve internal policy, draft replies, and recommend refunds below a defined amount.
HarborHelp is invented; the proposal is not. On 27 February 2024 Klarna announced an assistant built with OpenAI. One month after going live globally, it “has had 2.3 million conversations, two-thirds of Klarna’s customer service chats”. It was “doing the equivalent work of 700 full-time agents”. It handled refunds, returns, cancellations and disputes in 23 markets and more than 35 languages. Errand resolution fell from 11 minutes to under 2. Klarna also reported a 25% drop in repeat inquiries. Fourteen months later the same company reopened recruitment for human agents. On 8 May 2025 its co-founder and chief executive, Sebastian Siemiatkowski, told Bloomberg that “as cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality.” Both the launch numbers and the correction are evidence. An audit that reads only the first is an audit of a press release.
- Current baseline: agents search several policy pages, write replies manually, and escalate refunds to supervisors.
- Claimed benefit: faster first response, more consistent policy language, and fewer supervisor interruptions.
- Available evidence: three years of tickets, agent edits, resolution codes, refund records, and customer survey responses.
- Proposed workflow: the assistant drafts; agents approve replies; low-value refunds may later become automatic.
- Known concerns: outdated policy documents, sensitive account data, inconsistent historical labels, peak-season queues, and users in four languages.
- Executive request: approve a six-week pilot and define evidence required before any autonomous refund action.
Your task is not to praise or condemn HarborHelp; it is to make its assumptions, evidence, controls, and decision criteria explicit.
Steps
Build the audit workbook
Complete each page in order. Do not jump to model choice before the workflow and evidence are clear.
Page 4 is where audit workbooks most often record a control that does not exist. “Agents approve replies” is a design intention. The operational question is what approval looks like on the two-hundredth ticket of a shift. Medicine has measured the answer. Heleen van der Sijs, Jos Aarts, Arnold Vulto and Marc Berg surveyed the published studies of prescribing systems. Their review appeared in the Journal of the American Medical Informatics Association in March 2006. It is called “Overriding of Drug Safety Alerts in Computerized Physician Order Entry”. They found that “drug safety alerts are overridden by clinicians in 49% to 96% of cases”. Overriding is frequently justified, because many alerts do not apply to the patient in front of the prescriber. That is precisely why the number belongs on Page 4. A review step whose override rate nobody measures cannot be told apart from a review step that is not happening.
- 1
Page 1 — Purpose and baseline
Describe the user problem, current process, affected people, and measurable reason to change.
- 2
Page 2 — Task decomposition
Separate classification, retrieval, ranking, generation, policy checking, and proposed action.
- 3
Page 3 — Evidence map
Trace ticket data, labels, policy sources, permissions, timing, missingness, and language coverage.
- 4
Page 4 — Decision system
Define thresholds, human review, refund limits, fallbacks, appeals, logs, and ownership.
- 5
Page 5 — Failure register
List common errors, severe errors, affected groups, detectability, reversibility, and remedy.
- 6
Page 6 — Evaluation plan
Choose baselines, offline tests, pilot outcomes, slices, stress tests, and stop conditions.
- 7
Page 7 — Recommendation
State proceed, narrow, measure first, postpone, or reject, with conditions and unresolved questions.
Comparison
Four plausible proposals—and why none is automatically correct
A strong decision memo considers alternatives rather than treating deployment as a yes-or-no vote on AI itself.
Three of the four options rest on the same load-bearing assumption. It is that a human reviewer catches what the system gets wrong. That assumption has been measured. Kate Goddard, Abdul Roudsari and Jeremy Wyatt showed 26 UK NHS general practitioners 20 prescribing scenarios with pre-validated answers. Six of the scenarios carried deliberately incorrect advice. Their paper appeared in the International Journal of Medical Informatics in 2014. It is called “Automation bias: empirical results assessing influencing factors”. Accuracy rose from 50.38% before the advice to 58.27% after it. At the same time, in 5.2% of all cases a decision that had been correct before the system spoke was wrong after it. The authors record “a net improvement of 8%”. Assistance helped on balance, and created a class of error the unassisted workflow did not have. A drafting-only pilot is the cheapest place to find out how large that class is for HarborHelp.
Drafting-only pilot
Agents receive summaries and draft replies, while all actions remain manual.
- Lowest authority
- Good for measuring edit burden
- Still carries privacy and factual risks
- Useful first evidence stage
Retrieval-first redesign
Improve policy search and document ownership before enabling generation.
- Targets a root workflow problem
- Creates cleaner evidence
- May deliver value without generation
- Reduces outdated-source risk
Hybrid assisted workflow
Classification, retrieval, and drafting are combined with mandatory agent review and limited policy rules.
- Tests end-to-end value
- Requires queue and interface design
- Needs structured correction reasons
- Appropriate for a bounded pilot
Autonomous refunds
The system issues low-value refunds under policy without case approval.
- Higher speed and consequence
- Needs fraud and abuse controls
- Requires appeal and rollback
- Should follow stronger evidence
Figure
Visual
The decision pathway you should be able to defend
A capstone recommendation is strong when another reviewer can trace how evidence and values produced the conclusion.
The step from “Hypothesis” to “Pilot evidence” is not a formality. The base rate is published. Ron Kohavi, Alex Deng, Brian Frasca, Toby Walker, Ya Xu and Nils Pohlmann presented “Online controlled experiments at large scale” in 2013. All six were at Microsoft. They set out a tenet: “We are poor at assessing the value of ideas”. Behind it was a figure. “Only one third of the ideas tested at Microsoft improved the metric(s) they were designed to improve.” Kohavi put the same finding on the Bing engineering blog on 8 August 2013. “When ideas are evaluated objectively in a controlled experiment, less than a third move the metrics they were designed to improve.” Those were ideas teams believed in enough to build and ship into a test. A memo that lets the HarborHelp hypothesis stand in for the pilot result is betting against a base rate. One of the largest experimentation programmes in the industry measured that rate on itself.
- 1
Need
Agents spend measurable time searching fragmented policy and rewriting routine responses.
- 2
Hypothesis
Better retrieval and assisted drafting may reduce handling time without reducing answer quality.
- 3
Pilot evidence
Compare human-only and assisted workflows on quality, edit distance, time, escalations, complaints, and privacy incidents.
- 4
Risk controls
Use source citations, document ownership, access limits, agent approval, language slices, and non-AI fallback.
- 5
Decision gate
Proceed only if value exceeds review cost and severe failure rates remain within predefined bounds.
- 6
Expansion gate
Consider limited autonomous refunds only after separate policy, fraud, appeal, and monitoring evidence.
Key idea
Set a failure budget before the pilot creates momentum
Teams define success metrics and postpone the rejection criteria. It is then easy, and very common, to reinterpret a disappointing result as “promising.”
HarborHelp needs stopping conditions for privacy violations, unsupported policy claims, unequal language performance, review overload, and customer harm. A failure budget turns concern into an operational decision rule.
For some systems this stops being good practice and becomes law. Article 14(4) of the EU AI Act governs high-risk systems. It requires that such a system be supplied with real oversight powers. The people assigned to oversee it must be enabled “to decide, in any particular situation, not to use the high-risk AI system”. They must be able “to otherwise disregard, override or reverse the output of the high-risk AI system”. They must also be able “to intervene in the operation of the high-risk AI system”. And they must be able to “interrupt the system through a ‘stop’ button or a similar procedure”. The same paragraph names the human failure mode those controls have to survive. Overseers must be able “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)”. HarborHelp may or may not fall in scope. Those three requirements are still a usable shape for the stop conditions this section asks for. Who may stop it, how they stop it, and what stops them deferring to it.
A pilot without stop conditions is an adoption campaign, not an experiment.
Write the recommendation as a bounded argument
A defensible recommendation names the proposed scope, baseline, expected value, evidence quality, unresolved risks, controls, owners, and next decision date. It also states what would change the recommendation.
Avoid “AI is ready” or “AI is too risky.” Use a bounded claim instead. For example: “Pilot retrieval and drafting for English and Italian billing tickets, with mandatory agent approval and predefined failure controls.”
Specific scope makes disagreement productive because reviewers can challenge assumptions rather than slogans.
Steps
Your final deliverable
Produce a concise artifact that someone outside the course could use in a real review meeting.
“Top ten failure modes” is not a course invention. It is failure modes and effects analysis, specified internationally in IEC 60812:2018. That edition, 3.0, was published on 10 August 2018. It covers hardware, software and processes, including human action. Its history also carries a warning about how you rank. For decades FMEA prioritised by Risk Priority Number: severity multiplied by occurrence multiplied by detection. That lets a trivial, frequent, easily detected failure outrank a rare and catastrophic one. Severity 3 with occurrence 10 and detection 4 scores 120. Severity 10 with occurrence 4 and detection 2 scores only 80. The joint AIAG & VDA FMEA Handbook, first edition issued June 2019, therefore “replaced the RPN (Risk Priority Number) with AP (Action Priority)”. AP is a lookup that grades every failure mode High, Medium or Low. It weights severity above occurrence, and occurrence above detection. Rank your ten failure modes; do not multiply them.
- 1
One system diagram
Show inputs, task stages, model outputs, controls, human roles, actions, and feedback.
- 2
One-page evidence table
List each major claim, supporting evidence, baseline, missing test, and owner.
- 3
Top ten failure modes
Rank by severity, likelihood, detectability, reversibility, and affected group.
- 4
Pilot scorecard
Define quality, time, cost, adoption, equity, privacy, security, and incident measures.
- 5
Decision memo
Recommend proceed, narrow, measure first, postpone, or reject with explicit conditions.
- 6
Reflection note
Name one assumption you changed after tracing the full system instead of judging the model alone.
What you should now be able to do
You can define AI operationally, distinguish major methods and task families, trace a system beyond its model, question data and objectives, design human handoffs, inspect reliability, and recognize risk.
Most importantly, you can make a bounded decision under uncertainty. That skill is the foundation for every later path, whether you study optimization, computer vision, MLOps, governance, or generative systems.
AI literacy is complete when you can connect a technical output to evidence, workflow, consequence, and an accountable decision.
Key takeaways
- A complete AI audit begins with the user need and current baseline, not with model selection.
- Complex products should be decomposed into task stages with separate evidence, controls, and permissions.
- Alternatives such as retrieval-first redesign or drafting-only pilots may create stronger evidence before higher-authority automation.
- Pilot plans need rejection criteria and stop conditions as well as success metrics.
- A decision memo should state scope, evidence, controls, owners, unresolved risks, and what would change the recommendation.
- The core AI-literacy skill is connecting outputs to evidence, workflows, consequences, and accountable decisions.