AI literacy basics
Humans and AI: Assistance, Automation, and Oversight
Explore how AI changes human work, when oversight succeeds or fails, and how to design calibrated trust, meaningful control, and effective handoffs.
By the end you can
- Distinguish assistance, recommendation, delegated decision, and autonomous action
- Explain automation bias, under-trust, and calibrated trust
- Recognize when human review is meaningful rather than ceremonial
- Design clear roles, handoffs, feedback, and fallback for a human-AI workflow
Key idea
Putting a human “in the loop” solves nothing by itself
Human review can catch errors, add context, and preserve accountability. It can also be a ritual. Overloaded reviewers approve outputs they have no way to verify, and the box gets ticked.
Effective oversight requires five things: time, information, authority, skill, and a realistic option to disagree. Remove one. The human now absorbs the blame without controlling the system.
A human checkpoint is meaningful only when the reviewer can understand, challenge, and change the outcome.
Key idea
Those conditions are now in law
Those conditions have now been written into law rather than left to good intentions. Article 14 of Regulation (EU) 2024/1689 — the AI Act, dated 13 June 2024 and published in the Official Journal on 12 July 2024 — requires that a high-risk system be supplied so that the people assigned to oversee it are enabled “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)” and “to decide, in any particular situation, not to use the high-risk AI system or to otherwise disregard, override or reverse the output of the high-risk AI system”. For remote biometric identification, listed at point 1(a) of Annex III, Article 14(5) is more specific still: no action or decision may be taken on an identification “unless that identification has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority”. Two reviewers, named competence, named authority — the legislator is describing the same conditions this section does, and has evidently seen the ritual version.
Start with the division of labor
AI can support perception, memory, search, drafting, forecasting, and consistency. People contribute goals, contextual judgment, moral responsibility, lived experience, negotiation, and the ability to recognize when the task itself is wrong.
Good design allocates work according to those strengths and to the consequences of failure. It never asks which is better in the abstract, humans or AI. The question has no answer.
The design unit is a team of people and tools working under constraints.
Comparison
Four degrees of delegation
The same model can occupy different roles. Moving right increases speed and scale, but also increases the importance of safeguards and recoverability.
Assist
The system organizes information or drafts material without recommending a final choice.
- User remains the primary reasoner
- Easy to compare alternatives
- Errors are usually reversible
- Example: summarize a long ticket
Recommend
The system proposes a choice, priority, diagnosis, or next step.
- Can anchor human judgment
- Needs supporting evidence
- Reviewer retains authority
- Example: suggest an escalation queue
Decide under policy
The system selects an outcome within approved bands and exceptions.
- Requires explicit policy
- Needs appeals and audit records
- Higher consequence per error
- Example: approve a low-risk refund
Act autonomously
The system changes the environment without case-by-case approval.
- Speed and scale increase
- Fallback and containment are critical
- Best for bounded reversible actions
- Example: adjust a warehouse route
Case
The ladder an industry had to write down
Whole industries have had to write this ladder down and argue about the rungs. SAE International’s Recommended Practice J3016, issued in January 2014 and last revised in April 2021, “provides a taxonomy with detailed definitions for six levels of driving automation, ranging from no driving automation (Level 0) to full driving automation (Level 5)”, and the US road-safety regulator NHTSA uses the same six levels. The contested rung is Level 3, conditional driving automation, where the vehicle performs the entire driving task but the definition keeps a person in it: a “DDT fallback-ready user” who is “receptive to ADS-issued requests to intervene and to evident DDT performance-relevant system failures in the vehicle”. Note also what the standard refuses to say. Section 8.1 states that “this document is not a specification and imposes no requirements” and “imposes no requirements, nor confers or implies any judgment in terms of system performance”, so classifying a feature as Level 4 rather than Level 3 says nothing about whether it is safe. A delegation label describes who is supposed to act; it never certifies that the safeguards exist.
Example
Five ways human-AI handoffs fail
Collaboration failures often arise from role design rather than raw model accuracy.
- Anchoring: a reviewer sees the AI suggestion first and stops considering plausible alternatives.
- Information asymmetry: the system uses signals the reviewer cannot inspect, making disagreement difficult to justify.
- Queue pressure: review volume exceeds available time, so “human oversight” becomes rapid confirmation.
- Responsibility gap: operators can override the system but are punished when they do, so authority exists only on paper.
- Skill erosion: routine automation reduces practice, leaving people less prepared for rare failures when manual control returns.
Bainbridge named these failures before the software existed
Most of these were named before any of them involved software. Lisanne Bainbridge set them out for process control in “Ironies of Automation”, Automatica 19(6), 775–779 (1983). On skill erosion she is exact: “physical skills deteriorate when they are not used, particularly the refinements of gain and timing. This means that a formerly experienced operator who has been monitoring an automated process may now be an inexperienced one.” Her one-sentence statement of the whole problem opens the paper’s section on human-computer collaboration: “By taking away the easy parts of his task, automation can make the difficult parts of the human operator’s task more difficult.” And the irony she keeps for last is the one training budgets still miss: “it is the most successful automated systems, with rare need for manual intervention, which may need the greatest investment in human operator training.”
Visual
A review loop that can actually work
Oversight is strongest when the system routes uncertainty deliberately and learns from structured corrections without treating every click as truth.
- 1
Present the task and evidence
Show the input, relevant sources, model proposal, scope limits, and uncertainty cues.
- 2
Ask for an explicit decision
Require accept, edit, reject, defer, or escalate rather than passive exposure.
- 3
Capture the reason
Record why the reviewer changed the output, especially for recurring categories.
- 4
Protect disagreement
Give reviewers authority, time, and incentives to challenge the system when needed.
- 5
Analyze corrections
Separate model errors, policy errors, ambiguous cases, and workflow problems.
- 6
Improve the right component
Update data, interface, rules, staffing, training, or model only after diagnosis.
Analogy
A second pilot needs instruments and authority
Two pilots share a cockpit, and the collaboration works because roles, instruments, checklists, communication, and transfer of control are explicit.
AI is often called a copilot. The label is empty unless the human can see the relevant evidence and take control, and it breaks entirely when the AI has no stable world model, or when nobody can inspect how its proposal was produced.
It is worth knowing that the cockpit did not start out explicit either. The change has a date. NASA convened a NASA/industry workshop, “Resource Management on the Flight Deck”. It ran in San Francisco from 26 to 28 June 1979. The proceedings were published in March 1980, edited by Cooper, White and Lauber. The FAA’s own published history of the field records what the meeting concluded. It also records what it named. The research presented there “identified the human error aspects of the majority of air crashes as failures of interpersonal communications, decision making, and leadership”. And “at this meeting, the label Cockpit Resource Management (CRM) was applied to the process of training crews to reduce ‘pilot error’ by making better use of the human resources on the flightdeck.” The analogy is available to borrow for one reason. Someone spent the following decades building the training, the callouts and the handover protocol underneath it.
“Copilot” is a workflow claim, not a guarantee of safe partnership.
Steps
Design oversight around failure, not reassurance
Oversight should be strongest where errors are hard to detect, difficult to reverse, or unevenly distributed.
The last step is the one most often skipped. The cost of skipping it has been measured. Fenton and colleagues studied computer-aided detection for screening mammography. They covered 429,345 mammograms from 222,135 women at 43 facilities in three US states between 1998 and 2002. The result appeared in the New England Journal of Medicine in 2007. Read together with the radiologists, the software moved specificity from 90.2% to 87.2%. It cut the positive predictive value from 4.1% to 3.2%, and raised the biopsy rate by 19.7%. The increases in sensitivity and in the cancer detection rate were not statistically significant. Their conclusion is a statement about the team, not about the detector: “the use of computer-aided detection is associated with reduced accuracy of interpretation of screening mammograms.” A component can be good and the pairing still worse.
- 1
Classify the consequence
Estimate severity, reversibility, time pressure, and who bears the cost of error.
- 2
Match review to difficulty
Send ambiguous or high-impact cases to reviewers with relevant expertise.
- 3
Expose useful evidence
Provide sources, comparisons, uncertainty, and known limitations without overwhelming the reviewer.
- 4
Preserve a non-AI path
Allow manual completion, deferral, escalation, or service recovery when the system is unavailable or inappropriate.
- 5
Test the team
Evaluate human-plus-AI performance, not only model performance or reviewer accuracy in isolation.
Figure
Position
Most human oversight you will meet is not oversight
This course takes a side here, because the phrase is used to close arguments rather than open them. When a vendor or a compliance document says a human is in the loop, the claim is worth nothing until you know four things: how long the reviewer gets per case, what evidence appears in front of them, whether they can say no without justifying it upward, and what happens to the queue when they do. Article 14 requires that whoever is assigned to oversee a high-risk system be able to disregard, override or reverse its output, and be kept aware of their own tendency to over-rely on it. The instructive part is that requiring that in law was necessary.
Bainbridge described the trap in 1983, before any of this involved a model: the better the automation, the less practised the person, and the harder the intervention they are being kept there to perform. A checkpoint that fails those four tests is not a safeguard. It is somewhere to put the blame afterwards — and when you find one, the right move is to write that in the review rather than tick the box.
Ask how many seconds a reviewer gets per case. The answer usually ends the discussion.
The goal is calibrated trust, not maximum trust
Over-trust leads people to accept plausible errors. Under-trust wastes useful assistance and can push users toward unmonitored workarounds.
Calibrated trust means relying on the system when the evidence supports reliance, and checking or rejecting it when conditions change. Products can help. They set realistic expectations, show the relevant evidence, and fail in ways a person can understand.
Both halves of that have been observed in one experiment. Tschandl and colleagues tested AI-based decision support for skin cancer recognition against clinicians at every level of experience. They reported in Nature Medicine in 2020. Good quality AI-based support “improves diagnostic accuracy over that of either AI or physicians alone”. And “the least experienced clinicians gain the most from AI-based support”. The same abstract records the other half: “faulty AI can mislead the entire spectrum of clinicians, including experts”. Expertise did not protect the experts. What separates the two situations is not how confident the reader feels. It is whether the reader can tell which system is in front of them. That is a property of the evidence the product puts on the screen.
Trust should track demonstrated reliability in the current context, not the fluency or confidence of the interface.
Key takeaways
- Human oversight is effective only when reviewers have information, time, expertise, authority, and a genuine option to disagree.
- AI can assist, recommend, decide under policy, or act autonomously; each role creates different safeguards and accountability needs.
- Automation bias, queue pressure, responsibility gaps, and skill erosion can make human-AI teams worse than either component appears alone.
- Corrections should be diagnosed as model, data, policy, interface, or workflow problems before triggering updates.
- A non-AI fallback and explicit transfer of control make systems more resilient.
- The goal is calibrated trust that tracks evidence and context rather than maximum confidence in the system.