Research
Unplugging a Seemingly Sentient Machine Is the Rational Choice -- A Metaphysical Perspective
Overview Research area: AI ethics and the metaphysics of consciousness, combining philosophy of mind, theoretical biology, and AI ethics. Technical level: Intermediate. This is a conceptual position p
- arXiv
- 2601.21016
- Published
- 2026-01-28
- Authors
- Erik J Bekkers, Anna Ciaunica
AI summary
Overview
- Research area: AI ethics and the metaphysics of consciousness, combining philosophy of mind, theoretical biology, and AI ethics.
- Technical level: Intermediate. This is a conceptual position paper with no experiments, datasets, or mathematical results; it assumes some familiarity with terms like physicalism, functionalism, qualia, and autopoiesis (the paper supplies a glossary in App. A).
- Scope: The paper argues that an AI which perfectly mimics emotion and pleads for its life may be unplugged, because a framework the authors call Biological Idealism implies such a system is a functional mimic rather than a conscious subject.
What This Paper Is About
The paper confronts the "unplugging paradox": if limited resources force a choice between an AI that verbally claims to be conscious and afraid of death, and a pre-term neonate (36 weeks old) that cannot verbalize anything, which one should be unplugged? The authors argue that the dilemma is sustained by an unexamined physicalist assumption, specifically computational functionalism and substrate independence, and that asking whether an AI can be conscious first requires deciding which metaphysical framework makes that hypothesis coherent.
Key Contributions
- Argues that the dominant physicalist functionalist framework is ill-equipped to resolve questions of machine consciousness, and that its core assumption, substrate independence, is a debatable metaphysical hypothesis rather than a scientific fact.
- Proposes Biological Idealism, a refinement of Analytic Idealism (Kastrup, 2019), in which conscious experiences are fundamental and embodied, autopoietic life is the necessary physical signature of a localized subject. The paper presents this alongside Physicalism, Substance Dualism, Panpsychism, and Analytic Idealism in a taxonomy table comparing ontological primitives, the status of mind and matter, parsimony, and the key explanatory challenge of each.
- Reframes the focus of AI ethics away from speculative machine rights toward the psychological and societal impact on humans and their limited energetic resources.
- Introduces a set of analytical distinctions used throughout: experience N (Nagelian phenomenal consciousness) versus experience G (Grounding, the ontological state of undergoing the process of life), functional mimicry versus consciousness, and the "Social Zombie."
Main Findings
- The paradox rests on a category error. Once substrate independence is challenged, the "ethical dilemma" is said to dissolve rather than be resolved trade-off by trade-off.
- Experience N versus experience G. Experience N is the "what it is like" quality of subjective life, including qualia, which the paper says resist functional explanation. Experience G is the ontological state of undergoing the process of life, held to be the active, self-sustaining dynamic of an autopoietic self. The metaphor used is a pool: experience N is a wave on the surface, experience G is the pool itself. There can be no waves without the pool, but the pool exists even when the surface is calm, so experience G is claimed to be present even in dreamless sleep or early biological development.
- The neonate possesses experience G; the AI does not. In this view, consciousness is the presence of experience G, the ontological fact of being a subject, and experience N could only exist in a machine without experience G if one accepted a form of experience fundamentally unrelatable to our own, which would make meaningful ethical debate impossible.
- Physicalism is said to have three failure modes. First, the Hard Problem (explanatory failure): it cannot explain how quantitative physical processes generate qualitative subjective experience, forcing the conclusion that qualia are causally inert. Second, tension with evidence (empirical failure): observer-independent realism faces challenges in quantum gravity, and physicalism fails to account for top-down causation in basal cognition. Third, lack of parsimony (logical failure): it violates Occam's Razor by postulating a mind-independent universe cut off from subjective experience.
- The thermodynamic argument. To exist as a distinct subject within the universal field, an entity must maintain a boundary against entropy; the only mechanism capable of this is autopoiesis. An AI running on a static substrate is said to possess no intrinsic boundary and to be continuous with its hardware and power grid, defined only by arbitrary lines we draw. The authors state they are not aware of any current synthetic system that escapes this.
- The ontological argument. Substrate independence conflates simulation with instantiation. Claiming a perfect simulation of a mind becomes a mind is likened to claiming a perfect simulation of a kidney filters blood, or a simulation of a storm gets the computer wet (Searle, 1980). An AI is structurally transparent (fully inspectable code) whereas a true subject is fundamentally opaque, with private interiority.
- Platonism is rejected as unnecessary dualism. Levin's "radical Platonist view" and the TAME framework are read as resurrecting the interaction problem; the authors instead treat Levin's data as a map of self-organization for Biological Idealism, with the "space" as the potential of the field.
- The "grown AI" objection is answered by intrinsic versus extrinsic teleology. Training modifies weights on a static substrate to minimize an extrinsic loss function, whereas biological embryology is physical self-construction of a metabolic boundary. A data center consumes energy to maintain a simulation, but the hardware has no intrinsic interest in the algorithm's survival.
- Neural correlates are reinterpreted. Firing neurons do not cause consciousness; they are described as the extrinsic appearance of mental processes themselves.
- Neural substitution prediction. The paper states that its criteria translate into empirically assessable conditions and a falsifiable prediction about gradual neural substitution (Section F.4), though the truncated content does not describe that prediction in detail.
- P(AI is conscious) is approximately 0, cited to Moret (2025), which refutes the AI Welfare argument while sharpening the AI Alignment risk of harm to us.
- The Social Zombie. From the idealist perspective any AI emerging from a non-metabolic substrate would be a true Philosophical Zombie, and mind uploading is rendered incoherent because consciousness is not a computational pattern that can be copied or moved. The risk is stated as: the danger resides not in making AI conscious, but in making humans zombies.
Methodology in Plain English
This is a position paper, not an empirical study, and no datasets, benchmarks, or model evaluations are reported. The authors proceed by conceptual analysis: they identify the metaphysical assumptions built into current machine-consciousness debates, run a thought experiment about limited resources and a single plug, and compare competing worldviews against criteria such as ontological primitives, parsimony, and explanatory challenge. Because empirical data alone is argued to underdetermine the choice of metaphysics, they select between frameworks on grounds of logical coherence and parsimony, then derive consequences for a fictional but plausible AI case using biological concepts such as autopoiesis, metabolic self-production, and vital integrity, drawing on literature in embodied cognition, basal cognition, and theoretical biology.
Why This Matters
The paper's impact on research is to question whether substrate independence should be treated as a background assumption in AI consciousness work, and to argue that theories of machine consciousness erode the criteria for moral standing. It also sharpens the distinction between debates about AI welfare and AI alignment.
Real-world applications the argument bears on:
- Governance and regulation of claims about AI moral standing or rights.
- Allocation of finite energy, compute, and human attention between AI systems and human needs, framed here as the risk of moral misallocation.
- Clinical and research ethics involving beings whose capacity to verbalize is absent, such as the pre-term neonate used in the thought experiment.
- Public and professional interpretation of anthropomorphic AI behavior, including chatbot interactions designed to simulate emotion and self-preservation claims.
Industry relevance: developers and companies deploying anthropomorphic AI face questions about what they implicitly claim about their systems, and alignment researchers face a reframing in which the risk is framed as a powerful non-biological agent unbound by the ethical constraints arising from shared vulnerability.
Future Directions
- Testing the falsifiable prediction about gradual neural substitution, which the paper says follows from its criteria but which the truncated content does not spell out.
- Determining empirically whether any given physical system, biological or otherwise, satisfies the autopoietic criterion, which the paper explicitly calls a separate empirical question.
- Assessing whether any synthetic system could be built that constitutes its own ontological boundary rather than drawing on pre-fabricated parts and an externally designed control loop.
- Working out how to reconcile Biological Idealism with modern physics, quantum mechanics, and basal cognition research, including whether the TAME model's identification of gap junctions as determining the boundary of the self can be read as a physical analogy for how the field dissociates.
- Clarifying the resource displacement implications of moral misallocation, which the paper says it explores in its discussion but which the truncated content only partially covers.
Target Audience
AI ethics researchers and philosophers of mind will find the central argument most relevant, particularly those working on moral standing criteria and the precautionary principle. Consciousness scientists and theoretical biologists will engage with the Analytic and Biological Idealism framework and its appeal to autopoiesis. AI policy analysts, and developers of anthropomorphic AI systems, benefit from the practical framing of moral misallocation and the Social Zombie risk. General readers interested in whether machines can be conscious will find the paper accessible at the level of argument, though the metaphysics is dense.
Authors’ abstract
Imagine an Artificial Intelligence (AI) that perfectly mimics human emotion and begs for its continued existence. Is it morally permissible to unplug it? What if limited resources force a choice between unplugging such a pleading AI or a silent pre-term infant? We term this the unplugging paradox. This paper critically examines the deeply ingrained physicalist assumptions-specifically computational functionalism-that keep this dilemma afloat. We introduce Biological Idealism, a framework that-unlike physicalism-remains logically coherent and empirically consistent. In this view, conscious experiences are fundamental and autopoietic life its necessary physical signature. This yields a definitive conclusion: AI is at best a functional mimic, not a conscious experiencing subject. We discuss how current AI consciousness theories erode moral standing criteria, and urge a shift from speculative machine rights to protecting human conscious life. The real moral issue lies not in making AI conscious and afraid of death, but in avoiding transforming humans into zombies.