Skip to content
AI.info

Research

Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research

Overview Research area: Philosophy of science and philosophy of AI — specifically the conceptual foundations of AI research, benchmark design, and AI safety. Technical level: Intermediate. The paper c

Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
arXiv
2511.15282
Published
2025-11-19
Authors
Ninell Oldenburg, Ruchira Dhar, Anders Søgaard

AI summary

Overview

Research area: Philosophy of science and philosophy of AI — specifically the conceptual foundations of AI research, benchmark design, and AI safety.

Technical level: Intermediate. The paper contains almost no mathematics, but it assumes familiarity with contemporary debates about scaling laws, AGI, benchmarks, and alignment.

Scope: The paper argues that AI research sits on a spectrum between two implicit conceptions of intelligence — Intelligence Realism and Intelligence Pluralism — and shows how those implicit commitments shape methodology, interpretation of evidence, and risk assessment.

A note on the source material: the provided content is truncated mid-sentence in the "Alignment Approaches" section, so any later sections of the paper are not available here. The paper reports no quantitative empirical results — no dataset sizes, no model performance numbers, and no benchmark scores.

What This Paper Is About

AI researchers examining the same scaling curves, model architectures, and benchmark results reach opposite conclusions: some read them as signs of progress toward superintelligence, others as task-specific optimization that is far from human-like cognition. The paper argues this disagreement can be traced to an unstated difference in what researchers believe intelligence is. Its goal is to make those implicit beliefs explicit and give researchers a shared vocabulary for separating genuine empirical disputes (which more data could settle) from paradigmatic ones (which require philosophical examination).

Key Contributions

  1. A definition of two poles. The paper defines Intelligence Realism (intelligence is a single, universal capacity measurable across all systems) and Intelligence Pluralism (intelligence consists of diverse, context-dependent capacities that cannot be reduced to one universal measure), each specified through three claims — ontological, epistemological, and methodological — plus five distinguishing assumptions.

  2. A diagnostic rubric. Table 1 offers an eight-dimension rubric for locating a research program on the spectrum: Ontology, Benchmarks, Comparison Claims, Scaling Laws, Failure Attribution, AI Risk Focus, Architecture Preference, and Success Criteria — each with realist, intermediate/mixed, and pluralist indicators.

  3. A demonstration that the distinction structures AI research in three areas. The paper shows how realist and pluralist commitments generate different approaches to methodology (model selection, benchmark design, validation), different readings of the same empirical phenomena (scaling laws, emergent capabilities, failure analysis), and categorically different AI risk assessments and alignment strategies.

  4. A targeted synthesis, not a new distinction. The authors explicitly state their contribution is not inventing a new dichotomy but contextualizing long-standing debates from psychology, philosophy, and cognitive science within contemporary AI research, following Kuhn (1962) on paradigm-dependent observation and Hanson (1958) on the theory-ladenness of observation.

Main Findings

  • Realism rests on five assumptions. Algorithmic universality (intelligence follows discoverable universal mathematical principles), single optimality (an optimal intelligence algorithm exists), intelligence variance through implementation (different systems are approximations of the same algorithmic core), reducibility (advanced intelligence reduces to constitutive features), and commensurability (all intelligence can be compared along unified dimensions). Cited examples include Hutter's AIXI model, Schmidhuber's "optimal ordered problem solver," Commons and Ross (2008) on cross-species metrics, Goertzel (2014), Hernández-Orallo (2017), Chollet (2019) on ARC, and Legg and Hutter (2007), whose universal intelligence measure assigns weights w_μ to environments μ in the space of all environments E, aggregating them into a single score.

  • Pluralism denies all three realist claims and rests on five counter-assumptions: algorithmic diversity, multiple equilibria, emergence, irreducibility, and incommensurability. Cited examples include desert ants' path integration (Wehner et al. 2003), which the paper says would be catastrophically suboptimal in forest environments; Clark's nutcrackers' spatial memory for thousands of seed locations (Sherry et al. 1992); and octopus distributed intelligence with semi-autonomous processing in the arms (Godfrey-Smith 2016), which the paper says challenges the unit of analysis for realists. Supporting traditions cited include Gardner (1983), Gigerenzer and Goldstein (1996), Wilson (2002), Simon (1990), and Lieder and Griffiths (2020).

  • The paper presents both sides of each dispute. For the ants, realists would reply that both ants and humans implement approximate solutions to the same abstract problem — optimal path planning under uncertainty — while pluralists would call "path planning" an analyst's construction imposed retrospectively. For the octopus, realists would say the unit of analysis is the optimization process, not the physical substrate. For emergence, realists invoke the analogy that physics has universal laws despite emergent phenomena, which pluralists counter by saying the analogy assumes what is in dispute.

  • Pluralism faces a self-undermining problem. If every successfully adapted cognitive strategy counts as intelligent, the term "intelligence" loses discriminatory power and becomes analytically vacuous. The paper suggests other, relational properties become more analytically useful instead: danger for other agents, efficiency given the environment and agent features, and innovativeness in comparison with other agents.

  • The two views are a continuum, not a hard classification. Intermediate positions are explicitly allowed — for example, holding ontological realism while practicing methodological pluralism, or holding ontological pluralism while decomposing each form of intelligence into algorithms. Multi-task benchmarks illustrate the middle of the spectrum: they acknowledge domain distinctions but still aggregate performance into unified rankings.

  • Methodology diverges in three places. Model selection: realism favors unified architectures such as general-purpose large language models and transformers; pluralism favors specialized architectures, with the paper citing work on epistemic language (Ying et al. 2025), virtual bargaining (Levine et al. 2024), logical reasoning (Olausson et al. 2023), online goal inference (Zhi-Xuan et al. 2020), and knowledge inference in lie production (Tan et al. 2024), and Griffiths et al. (2024) on emulating a whole brain by combining specialized architectures. Benchmark design: realism favors aggregated benchmarks such as BIG-Bench (Kazemi et al. 2025) and ARC-AGI (Chollet 2019); pluralism favors separate domain evaluations such as ConceptARC (Moskvichev et al. 2023), FANToM (Kim et al. 2023), NormAd (Rao et al. 2024), and EWoK (Ivanova et al. 2024). Validation: realism judges success by progress toward universal principles and by equivalence or superiority to biological intelligence on standardized tasks; pluralism judges success by understanding specific cognitive phenomena in their ecological contexts.

  • The same data supports opposite interpretations. For scaling laws, Bubeck et al. (2023) read GPT-4 as "an early (yet still incomplete) version of an artificial general intelligence (AGI) system," while Mitchell (2021) cites Dreyfus (2012) and the analogy that "the first monkey that climbed a tree" was not making progress toward landing on the moon. For emergent capabilities, the paper cites Schaeffer et al. (2023) for the claim that apparent discontinuities are artifacts of benchmarks with arbitrary pass/fail thresholds — the emergence lives in the metrics, not the underlying process. For failure analysis, Bubeck et al. (2023) characterize limitations in long-term planning, temporal reasoning, and causal understanding as "missing components" fixable by further training, while pluralists cite systematic failures in common-sense physical reasoning (Ivanova et al. 2024), Theory of Mind (Kim et al. 2023), and non-human-like vision processing (Bowers et al. 2023) as evidence of fundamental architectural mismatches.

  • Risk assessment diverges categorically. Realists focus on superintelligence and unified alignment solutions, with the paper citing Bostrom (2014), Amodei et al. (2016), Constitutional AI (Bai et al. 2022), scalable oversight (Zeng et al. 2025; Bowman et al. 2022), and mesa-optimization and deceptive alignment (Hubinger et al. 2019). Pluralists prioritize context-specific, near-term harms — the paper names algorithmic bias in criminal justice, labor displacement (Gebru and Torres 2024), and misinformation (Bender 2024) — and argue current psychometric tests miss meta-learning, causal reasoning, innate priors, and robust generalization (Raji et al. 2021; Mitchell 2021). The paper notes the boundary is not clear-cut and that some programs, such as Hendrycks et al. (2023), acknowledge both kinds of risk.

Methodology in Plain English

This is a conceptual and philosophical paper, not an experimental one. The authors read and synthesize existing debates from AI research, psychology, philosophy, and cognitive science, then distill them into a two-pole spectrum with a set of defining assumptions. They build a diagnostic table of indicators that researchers can use to identify where a given research program sits. They then apply this framework to three domains — how research is done, how evidence is read, and how risk is assessed — illustrating each with published examples from the AI literature. The analytical lens comes from Kuhn's account of paradigms determining what counts as significant evidence and Hanson's account of observation being theory-laden. No experiments, datasets, or model evaluations were run.

Why This Matters

Impact on research. The framework gives AI researchers a vocabulary for identifying when a disagreement is empirical (solvable with more data) versus paradigmatic (requiring explicit philosophical examination). This matters because the field currently spends significant effort arguing past each other while looking at identical evidence.

Real-world applications the paper highlights (the paper does not present applied case studies; these are the domains it names as stakes):

  • Benchmark and evaluation design — whether to build aggregated benchmarks like BIG-Bench and ARC-AGI or separate domain evaluations like ConceptARC, FANToM, NormAd, and EWoK.
  • Model architecture and research investment — unified general-purpose models versus specialized or modular systems built around specific cognitive phenomena.
  • AI safety and governance — the paper notes that if a system cannot be assessed on metrics that make sense to humans, assessing whether it will become dangerous is a major open challenge for pluralists, while realists face the challenge of whether alignment properties transfer across capability levels.
  • Domain-specific societal harms — algorithmic bias in criminal justice, labor displacement, and misinformation, which pluralists treat as higher priority than hypothetical superintelligence scenarios.

Industry relevance. The methodological implications map directly onto commercial decisions: which architectures to fund, how to design and report evaluations, whether to invest in one general capability pipeline or many specialized ones, and how to allocate safety resources between near-term deployment harms and long-horizon speculative risks. The paper also argues the framework enables more productive policy debates by clarifying which disagreements are about values and paradigms rather than about facts.

Future Directions

  • Testing the framework empirically. The paper offers a conceptual rubric but does not report any study measuring whether researchers' stated positions on the realism–pluralism spectrum predict their methodological choices or interpretations.

  • Developing intermediate positions. The authors allow that researchers need not be consistent across ontological, epistemological, and methodological dimensions, but do not work out a systematic account of how mixed positions should be evaluated or how they resolve tensions.

  • Governance under pluralism. The paper explicitly raises a challenge it does not resolve: if pluralists are right that systems solve fundamentally different computational problems through categorically different mechanisms, and cannot be validated against human metrics, how can anyone assess whether such a system is dangerous?

  • Reducing the risk of a vacuous concept. The paper notes that under pluralism, "intelligence" risks becoming analytically empty, and proposes relational properties — danger, efficiency, innovativeness — as replacements, but leaves the systematic development of that alternative vocabulary to future work.

Target Audience

AI researchers who want to understand why their colleagues read the same results differently; philosophers of science and of AI working on foundations of intelligence and cognition; AI safety, alignment, and governance researchers deciding how to prioritize and frame risks; benchmark designers and evaluation researchers choosing between aggregated and domain-specific approaches; and policymakers and graduate students who need a compact map of the conceptual disagreements underlying current AI debates.

Authors’ abstract

In this paper, we argue that current AI research operates on a spectrum between two different underlying conceptions of intelligence: Intelligence Realism, which holds that intelligence represents a single, universal capacity measurable across all systems, and Intelligence Pluralism, which views intelligence as diverse, context-dependent capacities that cannot be reduced to a single universal measure. Through an analysis of current debates in AI research, we demonstrate how the conceptions remain largely implicit yet fundamentally shape how empirical evidence gets interpreted across a wide range of areas. These underlying views generate fundamentally different research approaches across three areas. Methodologically, they produce different approaches to model selection, benchmark design, and experimental validation. Interpretively, they lead to contradictory readings of the same empirical phenomena, from capability emergence to system limitations. Regarding AI risk, they generate categorically different assessments: realists view superintelligence as the primary risk and search for unified alignment solutions, while pluralists see diverse threats across different domains requiring context-specific solutions. We argue that making explicit these underlying assumptions can contribute to a clearer understanding of disagreements in AI research.

Read the original paper