Research
When AI Does the Work, What Is Learning For? Post-Instrumental Learning and the Risk of Capacity Dissolution
Overview Research area: AI ethics, AI governance, and the philosophy of learning and education (arXiv category: cs.CY, AI Safety & Ethics). Technical level: Beginner-Friendly. The paper contains no ma
- arXiv
- 2607.28041
- Published
- 2026-07-30
- Authors
- Kai Yao
AI summary
Overview
- Research area: AI ethics, AI governance, and the philosophy of learning and education (arXiv category: cs.CY, AI Safety & Ethics).
- Technical level: Beginner-Friendly. The paper contains no mathematics, no experiments, and no datasets; it is a normative conceptual argument that uses one thought experiment as a stress test.
- Scope: In one sentence: the paper argues that when AI can produce the essays, code, reports, plans, and decisions through which institutions recognize competence, the case for learning must shift from producing outputs to preserving five capacities (end-setting, reason-giving, contestability, refusal/revision, participation), whose erosion the author calls "capacity dissolution."
What This Paper Is About
Existing AI ethics focuses on present failures such as bias, opacity, hallucination, labor extraction, privacy risk, and weak accountability, but the paper argues that if the case for learning rests only on those failures, then every technical improvement weakens it. To test this, the author imagines "perfect AI" that executes specified tasks flawlessly while having no authority over purposes, legitimacy, and responsibility, and asks what people and institutions must still learn. The goal is to give a positive, governance-oriented answer called post-instrumental learning, illustrated mainly through educational assessment under generative AI.
Key Contributions
- A three-way distinction between kinds of obsolescence. The paper separates task obsolescence (most people no longer need to do a task), skill obsolescence (a technique becomes less worth teaching), and the stronger claim of formative obsolescence (the practice that cultivated a capacity no longer matters). It argues that policy debate often collapses these together.
- A positive account of post-instrumental learning built on five capacities. End-setting, reason-giving, contestability, refusal/revision, and participation are presented as a proposed minimum functional set, each tied to a specific institutional failure mode in Table 1.
- The concept of capacity dissolution. Defined as a breakdown in the relation among an output, the person or institution relying on it, and the practice that makes the claim answerable. It is distinguished from deskilling, automation complacency, and diminished contestability, though it may coexist with them.
- An application to assessment and to governance. The paper analyzes three assessment models (product-only, authorship enforcement, and post-instrumental assessment), proposes a practical accountable-use design sequence, and extends the argument to delegation in general, including "distributional audits" and the distinction between exclusionary and formative friction.
Main Findings
- Perfect execution does not settle legitimacy. By stipulation, "perfect AI" has no rightful authority to decide which purposes should govern, whose interests count, what dependencies are acceptable, or who must answer for results. Flawless means therefore leave the surrounding normative questions untouched.
- Four limits that perfect AI execution cannot remove: accuracy cannot choose ends (end-setting is separate from representation); safety is not legitimacy (a safe system may still lack standing or recourse); detection is not integrity (the issue is whether an artifact still evidences the capacity an institution certifies); and simulation is not participation (a perfect explanation does not give the experience of making a claim that others can challenge).
- The risk can arrive disguised as improvement. Cleaner essays, faster classifications, smoother services, and better-formatted reports may appear while the surrounding practice thins out; the loss typically shows up later, when a student cannot defend a claim or a caseworker cannot explain a decision.
- Capacity loss is cumulative and political. Because acceptable outputs keep arriving, pressure favors further delegation; each successful delegation can remove another occasion for learning how to govern delegation. Capacity is therefore treated as sustained by recurring occasions for use, not as a private stock that survives once acquired.
- Five capacities map to five failure modes (Table 1): end-setting fails when optimized means drift from legitimate purposes; reason-giving fails when explanations become display rather than accountability; contestability fails when affected people encounter decisions as settled facts; refusal/revision fails when dependence becomes compulsory even when harmful; participation fails when outputs remain but the practice that made them meaningful thins out.
- Assessment should begin with validity, not enforcement. A cheating-first policy "starts too late"; the prior question is what capacity the assessment is meant to evidence and whether the artifact still provides credible evidence of it. Evidence may need to move into oral defense, process commentary, source comparison, in-class transfer, or version histories.
- Three assessment models are compared: product-only (can certify the appearance of capacity without capacity), authorship enforcement (can make detectable human production the point, invites surveillance and uneven suspicion), and post-instrumental assessment (asks whether the learner can stand in an accountable relation to AI-mediated work). Post-instrumental assessment is described as stricter than permissive product assessment and less brittle than blanket prohibition.
- Friction is not automatically good. The paper distinguishes exclusionary friction (needless paperwork, inaccessible interfaces, opaque appeals) from formative friction (stating a goal before optimizing, comparing alternatives, confronting a counterargument, explaining a choice before high-stakes delegation). Formative friction should be proportionate, accessible, and challengeable.
- AI literacy is insufficient on its own. Competent users cannot repair systems that give them no role, and appeal processes have little value when people have not learned how to use them; the paper calls for capacities distributed across affected people, intermediaries, and communities.
- No empirical results are reported. The paper states explicitly that it is not a prediction that perfect AI is near, not an empirical measurement of learning loss, and not a design proposal for one educational technology.
Methodology in Plain English
The paper is a normative conceptual analysis rather than an experimental study. The author deliberately adopts a generous assumption, "perfect AI," meaning perfect task execution once goals, constraints, and roles have been specified, and uses it as a stress test: if the case for learning survives even after familiar AI failures are removed, then the case does not depend on those failures. The author then reasons through what delegation still requires under that assumption, drawing on automation research (which shows high system performance does not eliminate oversight, intervention, and responsibility allocation), on accountability and contestability literature, and on educational and philosophical work on practical judgment, participation, and technological mediation. The main worked case is formal educational assessment, though the author notes this does not define the paper's account of learning or knowledge, and says the standard is institutional as well as individual. A "practical right to understand" is offered as a governance standard, not a fully specified legal right.
Why This Matters
For research, the paper reframes the AI-and-learning debate away from whether AI can produce artifacts and toward which learned capacities must remain for delegated work to stay answerable. It positions AI literacy and AI education alongside governance debates over meaningful human control, contestability, reviewability, and recourse, and it adds a new object for governance review: the learning ecology around a system, not just its outputs.
Real-world applications:
- Education and assessment: redesigning assessment so that evidence tracks the learner's accountable relation to AI-mediated work, using brief justifications, defenses of contested premises, comparisons between generated and primary-source claims, vivas, or in-class transfer tasks rather than exhaustive prompt logs or blanket prohibition.
- Healthcare: preserving clinical judgment when risk scores, discharge summaries, and recommendations are generated, so clinicians can reconstruct and contest recommendations and notice clinically implausible numbers.
- Public administration and benefits decisions: ensuring affected people have usable routes to reasons, recourse, and revision rather than facing classifications as settled facts.
- Workplaces and professional formation: treating capacity as part of the system being maintained, giving newcomers hard cases and experienced workers time to review borderline outputs instead of becoming passive signers.
Industry relevance: the paper argues that procurement, review processes, training, and workflow design should be evaluated not only for accuracy, safety, and efficiency but for what repeated reliance teaches people not to do. A key claim is that "abdication" often appears as a sequence of individually reasonable local decisions, such as a teacher using a model to draft feedback because a class is too large, and that the resulting pattern is troubling only when no one asks which capacities will still be practiced, by whom, and under what conditions.
Future Directions
- Empirically measuring capacity dissolution. The paper defines the concept but explicitly does not measure learning loss; operationalizing and measuring it is left open.
- Building distributional audits into practice. The paper proposes that a distributional audit ask who becomes more capable and who becomes less capable through a system, meaning a structured institutional check rather than only a statistical compliance audit; how such audits should be conducted is not specified.
- Designing and testing post-instrumental assessment. The practical design sequence (name the target capacity, decide which AI uses scaffold versus replace it, require evidence where judgment matters) and its equity checks need concrete implementation and evaluation.
- Specifying the "practical right to understand." The paper offers this as a governance standard rather than a fully specified legal right, leaving its legal and institutional form unresolved.
Target Audience
This paper is most useful to educators, assessment designers, and university or school policy makers; to AI governance, policy, and ethics researchers working on accountability, contestability, and meaningful human control; to institutional decision makers in healthcare, public administration, and professional workplaces who are adopting AI systems; and to readers who want a conceptual argument about AI and learning that does not depend on current systems failing.
Authors’ abstract
As AI systems become capable of producing the essays, code, reports, summaries, plans, and decisions through which institutions usually recognize competence, a familiar question becomes harder to answer: what is learning for? Existing AI ethics rightly emphasizes present failures--bias, opacity, hallucination, labor extraction, privacy risk, and weak accountability. But if the case for learning rests only on those failures, then each technical improvement appears to weaken it. This article develops a different answer. Using the idealization of AI that executes specified tasks flawlessly while lacking authority over purposes, legitimacy, and responsibility, we argue for post-instrumental learning: learning that preserves the capacities people and institutions need when many useful outputs can be delegated. We analyze five such capacities--end-setting, reason-giving, contestability, refusal/revision, and participation--and name their erosion capacity dissolution. The central case is assessment under generative AI. When a polished artifact no longer reliably evidences understanding, institutions must assess the learner's accountable relation to AI-mediated work rather than the artifact alone. The takeaway is practical: AI governance should evaluate not only whether systems perform well, but also whether their deployment leaves people able to understand, challenge, revise, and share responsibility for the practices those systems mediate.