Skip to content
AI.info

AI literacy basics

Humans and AI: Assistance, Automation, and Oversight

Explore how AI changes human work, when oversight succeeds or fails, and how to design calibrated trust, meaningful control, and effective handoffs.

By the end you can

Key idea

Putting a human “in the loop” solves nothing by itself

Human review can catch errors, add context, and preserve accountability. It can also be a ritual. Overloaded reviewers approve outputs they have no way to verify, and the box gets ticked.

Effective oversight requires five things: time, information, authority, skill, and a realistic option to disagree. Remove one. The human now absorbs the blame without controlling the system.

A human checkpoint is meaningful only when the reviewer can understand, challenge, and change the outcome.

Key idea

Those conditions are now in law

Those conditions have now been written into law rather than left to good intentions. Article 14 of Regulation (EU) 2024/1689 — the AI Act, dated 13 June 2024 and published in the Official Journal on 12 July 2024 — requires that a high-risk system be supplied so that the people assigned to oversee it are enabled “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)” and “to decide, in any particular situation, not to use the high-risk AI system or to otherwise disregard, override or reverse the output of the high-risk AI system”. For remote biometric identification, listed at point 1(a) of Annex III, Article 14(5) is more specific still: no action or decision may be taken on an identification “unless that identification has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority”. Two reviewers, named competence, named authority — the legislator is describing the same conditions this section does, and has evidently seen the ritual version.

Start with the division of labor

AI can support perception, memory, search, drafting, forecasting, and consistency. People contribute goals, contextual judgment, moral responsibility, lived experience, negotiation, and the ability to recognize when the task itself is wrong.

Good design allocates work according to those strengths and to the consequences of failure. It never asks which is better in the abstract, humans or AI. The question has no answer.

The design unit is a team of people and tools working under constraints.

Comparison

Four degrees of delegation

The same model can occupy different roles. Moving right increases speed and scale, but also increases the importance of safeguards and recoverability.

FigureComparison · 4 columns

Assist

The system organizes information or drafts material without recommending a final choice.

  • User remains the primary reasoner
  • Easy to compare alternatives
  • Errors are usually reversible
  • Example: summarize a long ticket

Recommend

The system proposes a choice, priority, diagnosis, or next step.

  • Can anchor human judgment
  • Needs supporting evidence
  • Reviewer retains authority
  • Example: suggest an escalation queue

Decide under policy

The system selects an outcome within approved bands and exceptions.

  • Requires explicit policy
  • Needs appeals and audit records
  • Higher consequence per error
  • Example: approve a low-risk refund

Act autonomously

The system changes the environment without case-by-case approval.

  • Speed and scale increase
  • Fallback and containment are critical
  • Best for bounded reversible actions
  • Example: adjust a warehouse route

Case

The ladder an industry had to write down

Whole industries have had to write this ladder down and argue about the rungs. SAE International’s Recommended Practice J3016, issued in January 2014 and last revised in April 2021, “provides a taxonomy with detailed definitions for six levels of driving automation, ranging from no driving automation (Level 0) to full driving automation (Level 5)”, and the US road-safety regulator NHTSA uses the same six levels. The contested rung is Level 3, conditional driving automation, where the vehicle performs the entire driving task but the definition keeps a person in it: a “DDT fallback-ready user” who is “receptive to ADS-issued requests to intervene and to evident DDT performance-relevant system failures in the vehicle”. Note also what the standard refuses to say. Section 8.1 states that “this document is not a specification and imposes no requirements” and “imposes no requirements, nor confers or implies any judgment in terms of system performance”, so classifying a feature as Level 4 rather than Level 3 says nothing about whether it is safe. A delegation label describes who is supposed to act; it never certifies that the safeguards exist.

Example

Five ways human-AI handoffs fail

Collaboration failures often arise from role design rather than raw model accuracy.

  • Anchoring: a reviewer sees the AI suggestion first and stops considering plausible alternatives.
  • Information asymmetry: the system uses signals the reviewer cannot inspect, making disagreement difficult to justify.
  • Queue pressure: review volume exceeds available time, so “human oversight” becomes rapid confirmation.
  • Responsibility gap: operators can override the system but are punished when they do, so authority exists only on paper.
  • Skill erosion: routine automation reduces practice, leaving people less prepared for rare failures when manual control returns.

Bainbridge named these failures before the software existed

Most of these were named before any of them involved software. Lisanne Bainbridge set them out for process control in “Ironies of Automation”, Automatica 19(6), 775–779 (1983). On skill erosion she is exact: “physical skills deteriorate when they are not used, particularly the refinements of gain and timing. This means that a formerly experienced operator who has been monitoring an automated process may now be an inexperienced one.” Her one-sentence statement of the whole problem opens the paper’s section on human-computer collaboration: “By taking away the easy parts of his task, automation can make the difficult parts of the human operator’s task more difficult.” And the irony she keeps for last is the one training budgets still miss: “it is the most successful automated systems, with rare need for manual intervention, which may need the greatest investment in human operator training.”

Visual

A review loop that can actually work

Oversight is strongest when the system routes uncertainty deliberately and learns from structured corrections without treating every click as truth.

FigureProcess · 6 steps
  1. 1

    Present the task and evidence

    Show the input, relevant sources, model proposal, scope limits, and uncertainty cues.

  2. 2

    Ask for an explicit decision

    Require accept, edit, reject, defer, or escalate rather than passive exposure.

  3. 3

    Capture the reason

    Record why the reviewer changed the output, especially for recurring categories.

  4. 4

    Protect disagreement

    Give reviewers authority, time, and incentives to challenge the system when needed.

  5. 5

    Analyze corrections

    Separate model errors, policy errors, ambiguous cases, and workflow problems.

  6. 6

    Improve the right component

    Update data, interface, rules, staffing, training, or model only after diagnosis.

Analogy

A second pilot needs instruments and authority

Two pilots share a cockpit, and the collaboration works because roles, instruments, checklists, communication, and transfer of control are explicit.

AI is often called a copilot. The label is empty unless the human can see the relevant evidence and take control, and it breaks entirely when the AI has no stable world model, or when nobody can inspect how its proposal was produced.

It is worth knowing that the cockpit did not start out explicit either. The change has a date. NASA convened a NASA/industry workshop, “Resource Management on the Flight Deck”. It ran in San Francisco from 26 to 28 June 1979. The proceedings were published in March 1980, edited by Cooper, White and Lauber. The FAA’s own published history of the field records what the meeting concluded. It also records what it named. The research presented there “identified the human error aspects of the majority of air crashes as failures of interpersonal communications, decision making, and leadership”. And “at this meeting, the label Cockpit Resource Management (CRM) was applied to the process of training crews to reduce ‘pilot error’ by making better use of the human resources on the flightdeck.” The analogy is available to borrow for one reason. Someone spent the following decades building the training, the callouts and the handover protocol underneath it.

“Copilot” is a workflow claim, not a guarantee of safe partnership.

Steps

Design oversight around failure, not reassurance

Oversight should be strongest where errors are hard to detect, difficult to reverse, or unevenly distributed.

The last step is the one most often skipped. The cost of skipping it has been measured. Fenton and colleagues studied computer-aided detection for screening mammography. They covered 429,345 mammograms from 222,135 women at 43 facilities in three US states between 1998 and 2002. The result appeared in the New England Journal of Medicine in 2007. Read together with the radiologists, the software moved specificity from 90.2% to 87.2%. It cut the positive predictive value from 4.1% to 3.2%, and raised the biopsy rate by 19.7%. The increases in sensitivity and in the cancer detection rate were not statistically significant. Their conclusion is a statement about the team, not about the detector: “the use of computer-aided detection is associated with reduced accuracy of interpretation of screening mammograms.” A component can be good and the pairing still worse.

FigureProcess · 5 steps
  1. 1

    Classify the consequence

    Estimate severity, reversibility, time pressure, and who bears the cost of error.

  2. 2

    Match review to difficulty

    Send ambiguous or high-impact cases to reviewers with relevant expertise.

  3. 3

    Expose useful evidence

    Provide sources, comparisons, uncertainty, and known limitations without overwhelming the reviewer.

  4. 4

    Preserve a non-AI path

    Allow manual completion, deferral, escalation, or service recovery when the system is unavailable or inappropriate.

  5. 5

    Test the team

    Evaluate human-plus-AI performance, not only model performance or reviewer accuracy in isolation.

Figure

Three costs realised, the benefit unproven — the concrete meaning of a component being good and the pairing still worse. Fenton et al., NEJM 356(14), 2007; the relative fall in predictive value is derived from the two published rates.

Position

Most human oversight you will meet is not oversight

This course takes a side here, because the phrase is used to close arguments rather than open them. When a vendor or a compliance document says a human is in the loop, the claim is worth nothing until you know four things: how long the reviewer gets per case, what evidence appears in front of them, whether they can say no without justifying it upward, and what happens to the queue when they do. Article 14 requires that whoever is assigned to oversee a high-risk system be able to disregard, override or reverse its output, and be kept aware of their own tendency to over-rely on it. The instructive part is that requiring that in law was necessary.

Bainbridge described the trap in 1983, before any of this involved a model: the better the automation, the less practised the person, and the harder the intervention they are being kept there to perform. A checkpoint that fails those four tests is not a safeguard. It is somewhere to put the blame afterwards — and when you find one, the right move is to write that in the review rather than tick the box.

Ask how many seconds a reviewer gets per case. The answer usually ends the discussion.

The goal is calibrated trust, not maximum trust

Over-trust leads people to accept plausible errors. Under-trust wastes useful assistance and can push users toward unmonitored workarounds.

Calibrated trust means relying on the system when the evidence supports reliance, and checking or rejecting it when conditions change. Products can help. They set realistic expectations, show the relevant evidence, and fail in ways a person can understand.

Both halves of that have been observed in one experiment. Tschandl and colleagues tested AI-based decision support for skin cancer recognition against clinicians at every level of experience. They reported in Nature Medicine in 2020. Good quality AI-based support “improves diagnostic accuracy over that of either AI or physicians alone”. And “the least experienced clinicians gain the most from AI-based support”. The same abstract records the other half: “faulty AI can mislead the entire spectrum of clinicians, including experts”. Expertise did not protect the experts. What separates the two situations is not how confident the reader feels. It is whether the reader can tell which system is in front of them. That is a property of the evidence the product puts on the screen.

Trust should track demonstrated reliability in the current context, not the fluency or confidence of the interface.

Key takeaways