Research
What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education
Overview Research area: AI safety and ethics applied to education policy — specifically assessment validity, credential validity, and the governance of human–AI division of cognitive labor in higher e
- arXiv
- 2607.19988
- Published
- 2026-07-22
- Authors
- Kai Yao
AI summary
Overview
- Research area: AI safety and ethics applied to education policy — specifically assessment validity, credential validity, and the governance of human–AI division of cognitive labor in higher education.
- Technical level: Intermediate. The conceptual argument is accessible to non-specialists, but the paper leans on assessment-validity and automation-research literature, and the empirical audit uses a pre-specified codebook scored by four LLMs.
- Scope (1 sentence): The paper proposes "cognitive stewardship," a four-part framework linking learning claims, delegation boundaries, evidence standards, and safeguards, and then audits verified public generative AI assessment guidance from 30 universities to see whether that guidance connects permission rules to evidence and protection.
What This Paper Is About
Generative AI has made a basic premise of assessment fragile: that a submitted essay, program, or design can stand as evidence of the learner's competence. The paper argues the real question is not whether a student used AI, but which cognitive operations moved from the learner to the system (what it calls "educational delegation") and whether the remaining evidence still supports the credential the institution awards. The goal is to give universities a way to make the certification logic visible — what learners may delegate, what they must still demonstrate, and how fair evidence will be protected.
Key Contributions
-
A definition and framing of "educational delegation." The paper names the relation among tool use, task design, and the certified claim, and separates it from misconduct. It introduces five recurring delegation types — access support, feedback support, process support, substitution, and output-verification support — and stresses that these describe the role of assistance in a task, not moral or fixed tool labels: the same use can be access support in one task and substitution in another.
-
The cognitive stewardship framework. Four aligned elements: the learning claim being certified, the delegation boundary specifying what AI may do, the evidence standard showing what observable work, explanation, verification, or defense remains, and the safeguards protecting privacy, accessibility, proportionality, equity, and appeal. A policy is "under-specified" when any element is missing. The paper positions this as distinct from learner-facing AI literacy, evaluative judgment, and AI-use scales, because its unit of analysis is the warrant behind a course, program, or credential.
-
An empirical audit of 30 university policy packages. The corpus covers five English-speaking national systems: the United Kingdom (11 packages), Australia (5), New Zealand (2), Canada (5), and the United States (7). Constructs scored were learning claim (0–3), delegation boundary (0–4), evidence standard (0–4), and eight safeguards counted as indicators.
-
A scenario stress test and archetype analysis. Each policy was tested against access support, feedback support, process support, substitution, output-verification support, and a domain-specific programming-workflow case, producing 180 policy-by-scenario observations. Policies were then grouped into four archetypes by which element of the four-part relation was most visibly missing or unstable.
Main Findings
-
Boundary–evidence asymmetry. Most audited packages drew some line around AI use: 24 reached a delegation-boundary score of at least 2, and the mean boundary score was 2.47/4. The weaker element was the evidence standard, with a mean of 1.89/4. Twenty-two policies scored higher on boundary than evidence, 3 tied, and only 5 reversed the gap.
-
Safeguards are present but sparse. Packages contained a mean of 2.75 of 8 possible safeguards. The most visible categories were privacy at 67 percent; detection caution at 47 percent; accessibility at 42 percent; appeal or due process at 33 percent; equity of tool access at 32 percent; vendor governance at 29 percent; non-AI alternatives at 22 percent; workload/proportionality at 4 percent. The paper notes safeguards are not yet developing as a visible bundle.
-
Policies are clearest when AI use resembles cheating. Substitution was the most actionable scenario, averaging 2.49/4, with clear answers in 83 percent of packages and direct coverage in 93 percent; prohibition was the most common outcome for it (93 percent). The other five scenarios ranged from 1.17 to 2.32 on actionability: access support 1.93, feedback support 2.32, process support 2.25, output-verification support 1.17, and programming workflow 2.07.
-
Legitimate support is conditionally covered but evidence-thin. Feedback and process support were directly covered in 80 percent of packages, but evidence requirements and safeguards were lower. Access support showed safeguard guidance (55 percent) higher than evidence requirements (31 percent). Output-verification support was the thinnest scenario: only 23 percent of packages directly covered it, despite the paper arguing that the capacity to inspect, challenge, and correct AI output is one of the strongest reasons to teach with AI.
-
Validity links lag behind classification. For each scenario the paper reports a separate validity link field recording whether policy explains why an answer preserves the learning or credential claim: 22 percent for access support, 32 percent for feedback support, 30 percent for process support, 34 percent for substitution, 16 percent for output verification, and 27 percent for programming workflow. The paper reads this as policies being able to classify an AI use far more often than they explain what makes the resulting assessment evidence valid.
-
Four policy archetypes. Nine packages were higher-clarity cases, four were boundary-forward, eight had limited evidence visibility, and nine were partially connected. Higher clarity did not mean more permissiveness; it meant public text more often connected a permission or prohibition to evidence and safeguards.
-
Six anonymized profiles illustrate the spread. Scores (learning claim 0–3, boundary 0–4, evidence 0–4, safeguards 0–8) were: broadly connected 1.75/3.75/3.25/6.00; boundary-rich 2.00/3.75/2.50/5.75; evidence-rich 2.00/3.00/3.25/5.00; balanced visibility 2.00/3.00/3.00/4.75; boundary-evidence gap 1.75/3.25/1.25/0.00; limited public visibility 0.50/1.25/0.25/0.50.
-
Scoring uncertainty is treated as a sensitivity flag, not validation. Mean standard deviation across the four LLM coders was 0.57 for delegation-boundary scores, 0.53 for evidence-standard scores, and 0.42 for learning-claim scores. High-variation cases appeared in 13 to 22 of the 30 policy packages depending on the scenario.
-
The credential argument does not rest on model weaknesses. The paper explicitly rejects defenses of education built on hallucination, bias, weak reasoning, or unreliability, calling those an unstable foundation. Its limiting case is a safe, accurate, cheap, broadly capable system: even then, "a credential cannot certify a human capacity while allowing unrestricted delegation of that very capacity."
Methodology in Plain English
The work has two halves. The first is conceptual: the author reviews assessment-validity, feedback, automation, and AI-ethics literature, then defines educational delegation and builds the cognitive stewardship framework, including an illustrative table of delegation-centered assessment patterns across five claim types and two worked course examples (a first-year writing course and an introductory programming course).
The second half is a policy audit. The author assembled a purposive corpus of 30 public institutional policy packages — the official public sources through which one university tells students or instructors how generative AI may be used in assessed work. Inclusion required a publicly reachable official institutional source, explicit relevance to generative AI in assessment or student work, and enough substantive guidance to apply the rubric; error pages, news-only items without operational guidance, and sources that could not be independently verified were excluded. The corpus was built to compare visible policy designs, not to estimate worldwide prevalence.
Scoring used a pre-specified codebook — a written rubric defining the policy constructs, score levels, scenario tests, and source-evidence rules. A positive score required a quoted or locatable passage; unsupported higher scores were not counted. Four open-weight LLM models applied the same codebook to the same verified public text, and scores were averaged to reduce dependence on any single model's calibration. Between-model agreement and standard deviation are reported as sensitivity measures only. The author states plainly that the audit did not include an independently human-coded comparison set, so agreement statistics cannot rule out errors shared across models or establish how expert human coders would score the text; exact score levels are therefore treated as exploratory descriptions, and interpretation rests on recurring patterns across constructs, scenarios, and source-grounded excerpts. The codebook operationalizes the proposed framework, so the audit is described as a diagnostic application of that normative lens rather than a framework-neutral test of institutional quality — a low score means an element was not sufficiently visible in public text under the rubric, not that internal practice is weak. Completeness scores are used only to group policy profiles, not as an overall quality score, because the components have different scales.
Why This Matters
For research, the paper reframes AI-in-education governance away from detection and misconduct toward a validity problem: whether accumulated evidence still justifies an institutional claim about a learner. It supplies a testable construct set (claim, boundary, evidence, safeguard) and a scenario protocol that others could apply, extend, or contest, and it connects assessment-validity literature to emerging AI governance questions about surveillance, exclusion, and platform dependence.
Real-world applications:
- Institutional policy revision. The findings are framed as a revision guide: a boundary-forward package's next step is an evidence standard; a safeguard-light package's next step is protection and recourse; a package with limited evidence visibility needs a shared public vocabulary before it can fairly discipline students or support instructors.
- Assessment redesign at course level. The writing-course and programming-course examples show how evidence can be distributed across diagnostic work, source maps, feedback workshops, revision memos, code walkthroughs, and test interpretation rather than resting on a single final artifact.
- Risk-tiered governance. The paper proposes applying the framework proportionally: low-stakes formative work may need only clear AI-use statements, source maps, short reflection memos, or in-class checkpoints, while higher-stakes credentials may require secure demonstrations, sampled oral defenses, provenance notes, accessibility review, and independent appeals.
- Accessibility and equity review. Accessibility is treated as a condition of stewardship rather than an exception: AI or related tools may legitimately support screen reading, translation, speech-to-text, executive-function scaffolding, dyslexia support, anxiety scaffolding, or multilingual access, and the relevant test is whether assistance bypasses the learning claim or enables access to it.
Industry relevance: the paper notes that institutions procure tools, define legitimate help, collect disclosures, set penalties, and certify competence, and that vendor lock-in can make schools dependent on platforms they cannot audit. It also warns that surveillance harms may fall unevenly on racialized, disabled, low-income, international, and linguistically marginalized learners when automated accusations or behavioral monitoring are layered onto existing suspicion. The paper reports vendor/procurement governance appearing in 29 percent of audited packages, which is directly relevant to ed-tech procurement and platform accountability.
Future Directions
- Add human-coded comparison. The audit explicitly lacks an independently human-coded comparison set, so a natural next step is comparing LLM coder scores against expert human coders to establish whether the observed patterns hold under human judgment.
- Move from public text to practice. The audit measures public policy design, not classroom practice, institutional intention, or learning outcomes. Testing whether policies that score well on evidence and safeguards actually produce different assessment behavior or learner outcomes remains open.
- Extend beyond the purposive corpus. The corpus covers public English-language guidance from five higher-education systems and is described as a structured pilot rather than a representative survey; broader or non-English sampling would test the generality of the boundary–evidence gap.
- Build out output-verification assessment. Output-verification support was the thinnest scenario, with only 23 percent direct coverage, yet the paper argues it is one of the strongest reasons to teach with AI — leaving open how to design, teach, and assess verification as a capability in its own right.
- Study enforcement and equity effects. The paper notes that if evidence is absent, later enforcement cannot reconstruct learner contribution without suspicion, workload, or unequal burden, and that disclosure may make students more legible without making assessment fairer — questions that require empirical follow-up outside the current scope.
Target Audience
Higher-education policy makers, quality-assurance and accreditation staff, and assessment or curriculum designers who must decide what a grade or credential still certifies under AI mediation. It is also relevant to AI-ethics and AI-safety researchers interested in human–AI division of labor and credential validity, to educators in professional and clinical programs where delegated decisions affect third parties, to disability and access specialists evaluating whether AI rules help or exclude learners, and to ed-tech procurement and vendor-governance staff concerned with lock-in, privacy, and monitoring. Readers seeking quantitative effect sizes on learning outcomes will not find them here — the paper reports policy-text diagnostics, not learner results.
Authors’ abstract
Generative AI is changing a basic premise of educational assessment: that submitted work can reliably evidence the human capacities a credential claims to certify. The challenge is not simply whether students use AI, but what remains inferable about learning when some cognitive work has been delegated to a system. This paper develops cognitive stewardship, a framework for AI-mediated assessment that links the learning claim, delegation boundary, evidence standard, and safeguards. We then audit verified public generative AI assessment guidance from 30 universities. Using a pre-specified scoring codebook--a written, source-grounded rubric--four open-weight LLM models applied the rubric as structured coders, with scores averaged to reduce dependence on any single model's bias. The audit shows that public policies are becoming better at classifying AI use than at explaining what evidence and protections preserve credential validity. Boundaries are more visible than evidence standards; safeguards are uneven; and guidance is clearest when AI use resembles final-output substitution rather than feedback, access, verification, or professional workflow. The takeaway is that permission categories are necessary but insufficient. Universities need policies that make the certification logic visible: what learners may delegate, what they must still demonstrate, and how institutions will protect fair evidence rather than merely monitor AI use.