Responsible AI
AI Literacy, Competence, and Role-Based Training
Design role-specific AI literacy, competence standards, and training evidence for builders, operators, leaders, reviewers, and affected staff.
By the end you can
- Explain why AI literacy must be role-specific, assessed through realistic performance, and refreshed as systems and responsibilities change
- Distinguish Awareness, Role competence, and Expert qualification
- Identify evidence that connects executives and board to procurement, legal, and audit
- Design a review that moves from map roles and decisions to refresh and monitor
Visual
Five roles, and none needs all of it
Boards, product owners, technical teams, operators and procurement each need a different competence. None of them needs all of it. The map below is not a hierarchy. A procurement officer who cannot ask a vendor whether a model has ever been externally validated fails one way. A clinician who cannot judge what a score is worth at the bedside fails another. One syllabus written for both misses each of them.
- 1
Executives and board
Risk appetite, oversight questions, accountability, and escalation.
- 2
Product and domain owners
Purpose, workflow, harms, acceptance criteria, and recourse.
- 3
Technical teams
Data, model, evaluation, robustness, privacy, security, and monitoring.
- 4
Operators and reviewers
Interpretation, workload, override, uncertainty, documentation, and incidents.
- 5
Procurement, legal, and audit
Vendor evidence, contracts, obligations, control testing, and assurance.
Example
A sepsis model with an AUC of 0.63, read at the bedside
The Epic Sepsis Model was a proprietary tool implemented at hundreds of US hospitals. Nobody outside the vendor had ever measured it. In 2021 Wong and colleagues ran the external validation the purchase never had: 38,455 hospitalizations of 27,697 adults at Michigan Medicine, between 6 December 2018 and 20 October 2019. JAMA Internal Medicine gave the result in one line: “The ESM had a hospitalization-level area under the receiver operating characteristic curve of 0.63 (95% CI, 0.62-0.64).” The model missed 1,709 of 2,552 sepsis patients (67%). It still fired alerts on 6,971 of 38,455 hospitalizations (18%). Every clinician on those wards was being asked to read that score at the bedside. Whoever signed for the tool had bought a model nobody had checked. Those are two different jobs. No general awareness module tells them apart.
- Uniform training: One syllabus teaches the vocabulary of a sepsis score to the clinician, the procurement officer and the incident responder alike. What each of them has to do with that score is not the same job.
- Role mismatch: The clinician needs to know what an area under the curve of 0.63 (95% CI, 0.62-0.64) is worth when the alert fires at 3 a.m. Procurement needed a different question, and needed it before signing: why had a model implemented at hundreds of US hospitals never been externally validated?
- Competence gap: Completion records show attendance. They do not show that anyone could have named 1,709 of 2,552 missed sepsis patients as the reason never to read the model's silence as reassurance.
- Operational risk: Alerts on 6,971 of 38,455 hospitalizations (18%) are a workload before they are a warning. The people carrying that workload need an override path and a named escalation, not a definition of machine learning.
- Governance illusion: A training log proves that training occurred, which was never in doubt. What established what this tool was actually worth was an external validation in JAMA Internal Medicine. No completion rate substitutes for evidence of that kind.
Builder, operator, leader: three different syllabuses
AI literacy means grasping what an AI system can do, where it stops working, what it puts at risk, and when it fits the job in front of you. Competence adds the skill to perform a role reliably. Training should be proportional to the person's authority, exposure and task. A builder handles the data, evaluates the model, keeps it secure, tests it for fairness and writes down what they did. A leader sets the risk appetite, names who answers for a decision, knows what to ask a vendor before signing, and reads the evidence. Affected staff need clear rights and safe reporting channels.
The operator's syllabus is not a matter of taste. Article 14(4) of Regulation (EU) 2024/1689 enumerates it in five points. Whoever holds human oversight must be enabled to understand the system's capacities and limitations and detect anomalies (a), to remain aware of automation bias (b), to correctly interpret the output (c), to decide not to use the system or to “disregard, override or reverse the output” (d), and to intervene or stop the system (e). Point (b) is worth reading as the legislator wrote it: “to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias), in particular for high-risk AI systems used to provide information or recommendations for decisions to be taken by natural persons”. Interpret, judge the pull of the machine, override, stop. That sequence is a legal specification, not a training designer's preference.
One syllabus lands differently on different roles. That is measurable, and it has been measured. Give 138 radiologists and 127 internal and emergency medicine physicians the same eight chest X-ray cases, with the same advice attached, some of it deliberately wrong. The two groups do not respond alike. Gaube and colleagues reported the split in npj Digital Medicine in 2021: “We define clinical susceptibility as the propensity to follow incorrect advice, and we find that 41.73% of IM/EM physicians are susceptible, i.e., they always give the wrong diagnosis with inaccurate advice. This is true only for 27.54% of radiologists.” In the other direction, 28.26% of radiologists refuted every piece of incorrect advice, against 17.32% of the IM/EM physicians. The abstract names the mechanism: “As a group, radiologists rated advice as lower quality when it appeared to come from an AI system; physicians with less task-expertise did not.” Same cases, same advice, two roles. Different competence at rejecting a bad output.
Same eight cases, same advice: 41.73% of the IM/EM physicians always gave the wrong diagnosis when the advice was inaccurate, against 27.54% of the radiologists. One syllabus cannot be right for both.
Steps
Test what people decide, not what they attended
Training that ends at a slide deck gets measured by counting who turned up. The five steps below measure what people decide instead. The first two are already written down in a public framework, with clause numbers you can cite. NIST's AI Risk Management Framework, published in January 2023, puts training under GOVERN 2: teams and individuals are “empowered, responsible, and trained for mapping, measuring, and managing AI risks”. GOVERN 2.1 requires roles, responsibilities and lines of communication to be documented and clear throughout the organization. That is step 1, mapping roles to the decisions they own. GOVERN 2.2 requires that “The organization’s personnel and partners receive AI risk management training to enable them to perform their duties and responsibilities consistent with related policies, procedures, and agreements.” Training tied to the duties a person must actually perform is step 2. Neither clause counts attendance. And the word “partners” is why the agency-supplied clerk is inside the scope too.
1. Map roles and decisions
Connect each role to the AI decisions it makes or influences.
2. Define competencies
Specify knowledge, skills, judgment, and escalation behavior.
3. Build realistic practice
Use scenarios, counterexamples, and the actual tools people will operate.
4. Assess performance
Test decisions and actions rather than only recall of policy language.
5. Refresh and monitor
Retrain after model, use, law, incident, or role changes.
Comparison
Awareness, Role competence, or Expert qualification?
Awareness recognizes the vocabulary. Role competence performs a task under realistic pressure. Expert qualification carries independent assurance. The line between the first two tiers has legal weight wherever the system is high-risk. Article 26(2) of Regulation (EU) 2024/1689 requires a deployer to assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support. A person who can recite the policy and name the vendor's terms, and who has never handled a live output or tested a control, sits at the first tier and not the second. Whatever the completion log records, and whatever the job title says.
Awareness
Recognizes basic terms and organizational policy.
- Appropriate for broad workforce orientation
- Does not demonstrate role performance
- Can use short refreshers and scenarios
- Evidence: comprehension check
Role competence
Performs a defined task under realistic conditions.
- Requires practice and observable assessment
- Includes escalation and boundary cases
- Must be refreshed after system changes
- Evidence: scenario or supervised performance
Expert qualification
Provides specialist judgment or independent assurance.
- Requires deeper technical or legal knowledge
- May need credentials and experience
- Must disclose conflicts and limitations
- Evidence: reviewed work and professional standards
Example
Competence measured as behavior, not attendance
The last drill is the awkward one: confirm that workload, interface and authority actually let the trained behavior happen. The documented version of that failure is a crash report. A vehicle running an automated driving system crashed in Tempe, Arizona on 18 March 2018. A trained safety operator was behind the wheel. The National Transportation Safety Board adopted its probable cause on 19 November 2019: “The National Transportation Safety Board determines that the probable cause of the crash in Tempe, Arizona, was the failure of the vehicle operator to monitor the driving environment and the operation of the automated driving system because she was visually distracted throughout the trip by her personal cell phone. Contributing to the crash were the Uber Advanced Technologies Group’s (1) inadequate safety risk assessment procedures, (2) ineffective oversight of vehicle operators, and (3) lack of adequate mechanisms for addressing operators’ automation complacency—all a consequence of its inadequate safety culture.” Reporting on the determination records that she “looked down 23 different times” in the final three minutes. The report also records what had changed around her. During September–October 2017 Uber ATG folded the work of two vehicle operators into one, noting that “the consolidation of responsibilities also increased the task demands on the now-sole operator”. Two of the three contributing factors the board named are environment, not syllabus.
- Competency matrix: Define what each role must know, do, and escalate. For an operator of a high-risk system the five capabilities enumerated in Article 14(4) — anomalies, automation bias, interpretation, override, stop — are the floor, not the ambition.
- Scenario test: Create one plausible failure the role must identify and manage: a score with discrimination of 0.63 (95% CI, 0.62-0.64) firing on a patient who is not septic, and staying silent on one who is.
- Evidence record: Store assessment result, date, system version, and remediation plan. The European Commission is explicit that this is enough: “There is no need for a certificate. Organisations can keep an internal record of trainings and/or other guiding initiatives.”
- Environment check: Confirm that workload, interface, and authority allow trained behavior in practice. The NTSB named “ineffective oversight of vehicle operators” and “lack of adequate mechanisms for addressing operators’ automation complacency” among the causes, alongside a consolidation that “increased the task demands on the now-sole operator”. No further module reaches any of the three.
Key idea
What Article 4 required, and what it asks for since 27 July 2026
Completing the training tells the board who sat through it. It does not tell them who is competent. A person who cannot recognize when a system is being used outside the boundary it was approved for, and does not know who to escalate that to, remains an operational risk despite finishing every module. No training compensates for impossible workload, missing authority, or incentives that punish caution. A trained person only acts on the training when the controls are usable and their managers back them.
The law most often quoted here has changed, and the change is recent enough that older material still misstates it. Article 4 of Regulation (EU) 2024/1689 applied from 2 February 2025. It required providers and deployers to “ensure, to their best extent, a sufficient level of AI literacy”. On 27 July 2026 that text was replaced in full by the Digital Omnibus on AI, Regulation (EU) 2026/1744, adopted 8 July 2026 and published in the Official Journal on 24 July 2026. The provision now reads: “Providers and deployers of AI systems shall take measures to support the development of AI literacy of their staff and other persons dealing with the operation and use of AI systems on their behalf, taking into account their technical knowledge, experience, education and training and the context the AI systems are to be used in, and considering the persons or groups of persons on whom the AI systems are to be used. This obligation does not require providers or deployers to guarantee any specific level of AI literacy of any individual.” Recital 8 gives the reason: the stringent obligations had created “an additional compliance burden, particularly for smaller enterprises”. The Commission's own announcement says the same thing in a sentence: “Previous AI literacy requirement for companies is simplified, with the Commission and the Member States taking a stronger role in promoting AI literacy”. Ensure a sufficient level became support the development. The guarantee was disclaimed outright.
The regulator had already said what the duty does not contain. “Article 4 of the AI Act does not entail an obligation to measure the knowledge of AI of employees.” That is the European Commission's own AI Literacy Questions & Answers, published 7 May 2025 and last updated 27 July 2026. It adds that “There is no one size fit all when it comes to AI literacy and no strict requirements or mandatory trainings are imposed.” Asked how organisations document what they have done, the AI Office answers: “There is no need for a certificate. Organisations can keep an internal record of trainings and/or other guiding initiatives.” No training format is mandatory. No governance structure is mandated. “Other persons” reaches contractors, service providers and clients, so the agency clerk operating the tool on your behalf is inside the sentence. Supervision and enforcement by national market surveillance authorities start from 2 August 2026.
What did not soften is the duty owed by a deployer of a high-risk system. The Digital Omnibus on AI left Article 26(2) untouched: “Deployers shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support.” The Commission's Q&A confirms that it survived the Article 4 amendment — “for those who deploy high-risk AI systems, the obligation to ensure that their staff is trained to ensure human oversight remains in place.” Competence, training, authority and support sit in a single sentence of the statute. That is this lesson's argument about environment, written as law. An organisation that trains a person and then withholds the authority or the support has not discharged Article 26(2), however complete its attendance records are.
Since 27 July 2026 Article 4 asks providers and deployers to “take measures to support the development of AI literacy”, and adds that it “does not require providers or deployers to guarantee any specific level of AI literacy of any individual”. Article 26(2), untouched, still demands competence, training, authority and support for whoever oversees a high-risk system.
Carry this AI literacy and competence boundary forward
Competence decays with every model update. The legal floor moves too. The operative wording of Article 4 changed on 27 July 2026, so a training record written against the old “ensure, to their best extent, a sufficient level of AI literacy” formulation is dated in a way its author cannot see from the inside. And a training record carrying no date beside a system version proves nothing about anything.
Define when a literacy and competence finding requires the reviewer to redesign, restrict, remedy, or retire the system. And record, beside every assessment, three things: the date, the system version it was taken against, and whether the person assessed actually holds the authority and the support that Article 26(2) of Regulation (EU) 2024/1689 places in the same sentence as their training.
Key takeaways
- AI literacy concerns capabilities, limits, risks, and appropriate use in context. Since 27 July 2026 Article 4 of Regulation (EU) 2024/1689, as replaced by the Digital Omnibus on AI, asks providers and deployers to “take measures to support the development of AI literacy” rather than to ensure, to their best extent, a sufficient level of it.
- Competence is role-specific, and it has to be shown by performance. Given the same eight chest X-ray cases and the same advice, 41.73% of the IM/EM physicians always gave the wrong diagnosis when the advice was inaccurate, against 27.54% of the radiologists (Gaube and colleagues, npj Digital Medicine, 2021).
- Builders, operators, leaders, procurement staff and auditors need different learning objectives. Article 14(4) of Regulation (EU) 2024/1689 spells out five for the operator alone, from awareness of automation bias to the power to “disregard, override or reverse the output”.
- Completion rates do not prove anyone can manage a real boundary case. The European Commission's AI Literacy Q&A says “There is no need for a certificate.”, and NIST's AI Risk Management Framework asks instead that personnel and partners “receive AI risk management training to enable them to perform their duties and responsibilities”.
- Training cannot repair missing authority, excessive workload, or harmful incentives. The NTSB named “ineffective oversight of vehicle operators” among the causes of the Tempe crash of 18 March 2018, and recorded a consolidation of two vehicle operators into one that “increased the task demands on the now-sole operator”.
- Refresh competence after meaningful system, role, legal, or incident changes. The Article 4 rewrite of 27 July 2026 is a legal change; the external validation Wong and colleagues published in JAMA Internal Medicine is new information about a system already in use; an undated record cannot show which assessments either of them invalidated.