Skip to content
AI.info

Research

On the Measure of a Model: From Intelligence to Generality

Overview Research area: AI evaluation methodology and the philosophy of AI measurement, combining conceptual analysis with formal learning theory (PAC bounds, multitask learning). Technical level: Int

On the Measure of a Model: From Intelligence to Generality
arXiv
2511.11773
Published
2025-11-14
Authors
Ruchira Dhar, Ninell Oldenburg, Anders Soegaard

AI summary

Overview

Research area: AI evaluation methodology and the philosophy of AI measurement, combining conceptual analysis with formal learning theory (PAC bounds, multitask learning).

Technical level: Intermediate. The argument is stated in accessible prose, but two sections rely on formal definitions, agent-characteristic curves, and PAC-style generalization bounds that require some machine learning theory background.

Scope: The paper argues that AI evaluation should be grounded in generality — measured breadth and reliability of performance across a task distribution — rather than in intelligence, and supports this with a conceptual analysis of three implicit assumptions plus two formal generalization results.

What This Paper Is About

Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are commonly treated as measures of LLM "intelligence," yet the paper claims these scores fail to predict the things practitioners care about, such as human preference and performance on real-world tasks. The authors dissect the implicit assumptions behind intelligence-focused evaluation and argue that only one of them — generality — survives scrutiny. They then reframe evaluation as a multitask learning problem that ties assessment directly to measurable performance breadth and reliability.

Key Contributions

  1. Identification of three assumptions behind intelligence-based evaluation. The paper names generality (we want general-purpose systems), stability (there is a fixed set of tasks worth mastering), and realism (intelligence is a real, general capacity), and traces how each appears in current AI research such as Big Bench, HELM, instruction-tuned models like T0, FLAN-T5, and OPT-IML, and IQ-style comparisons.

  2. A demonstration that generality is independent of the other two assumptions. Through a formal setup defining generality as expected performance under a task distribution, stability as aggregated performance over a fixed subset, and realism as performance derived from a latent cognitive vector, plus a "Three Robot Designers" thought experiment using handball, badminton, and polo, the authors argue generality can be pursued without fixing a task set or positing a latent variable.

  3. An argument that only generality is necessary for coherent evaluation. Using the agent-characteristic curve of Hernández-Orallo et al. (2021), the paper shows that any evaluation rule depending only on mean performance is non-identifiable, and that the spread of the curve — the reciprocal quantity Γ_M = 1/S_M — is what makes evaluation stable under distributional shift.

  4. Two formal results grounding generality in multitask learning. Theorem 1 and Theorem 2 show a generalization bound reduction of approximately a factor of √n from evaluating across n tasks, with the error split into within-task and across-task components.

Main Findings

  • Intelligence benchmarks do not track what matters. The paper reports that performance trends on ARC-AGI-1, ARC-AGI-2, and LMArena differ significantly, so a model scoring higher on ARC is not necessarily better on LMArena. The figures show limited correlation between intelligence benchmarks and human preference or task performance; the paper does not report specific numeric scores or correlation values for these comparisons.

  • Task-specific benchmarks also diverge from benchmark intelligence. Figure 2 compares model performance on OpenBookQA, Entity Extraction, and StackUnseen, and the authors state these trends do not translate into universal capability on real-world tasks. Again, no numeric values are reported in the paper content.

  • Intelligence is conceptually unstable. The paper traces the dispute to Spearman and Thomson, notes continued disagreement in neuroscience, cognitive science, and education, and cites research on distributed, task-specific neural activation suggesting intelligence is not a unitary system but an emergent property of reorganizing task-specific systems.

  • Realism and stability do not survive the formal analysis. Fixing a benchmark subset removes the difficulty structure needed for comparison, and positing a latent "intelligence" adds nothing to the observable shape of the performance-versus-difficulty curve. The paper's bound analysis also notes that the three robot designers face incompatible task gradients, which no finite latent vector captures.

  • Multitask evaluation gives tighter guarantees. Combining within-task PAC bounds with across-task concentration via Hoeffding's inequality yields an overall bound of the form O(√((C + ln(n/δ))/m) + √((C + ln(1/δ))/n)). The within-task term shrinks with the number of samples per task (m), and the across-task term shrinks with the number of sampled tasks (n), producing the approximately √n reduction.

  • Three common objections are addressed. The authors respond to "we are already doing multitask learning" (training mixes tasks, but evaluation does not use task distributions or generalization bounds), "we don't need generality" (narrow systems are brittle and hard to maintain as tasks drift), and "intelligence is real" (correlations among task performances can be explained by shared inductive biases, so Occam's razor removes the need for the construct).

Methodology in Plain English

The paper does not train models or run new leaderboard experiments. Instead it combines three kinds of analysis.

First, a conceptual audit: the authors catalogue how benchmarks such as ARC, Raven tests, the Blackbird Task, Big Bench, and HELM are used, and separate the assumptions built into that usage.

Second, a formal setup: a task environment is modeled as a probability measure Q over a (possibly infinite) set of tasks, with each model inducing a performance function that maps a task to a normalized score in [0,1]. Generality is the expectation of that function under Q, stability is an aggregation over a fixed subset S, and realism adds a latent vector plus task-specific decoders. Tasks are then ordered by difficulty, and performance is aggregated at each difficulty level into an agent-characteristic curve, whose spread S_M (and its reciprocal Γ_M) serves as the measure of generality.

Third, a proof exercise: the authors extend classical PAC-learning bounds to the multitask setting by splitting error into a within-task component and an across-task component, then combine them with the triangle inequality (detailed in Appendix A). An illustrative thought experiment about three robot designers choosing different evaluation assumptions concretizes the formal distinction.

Why This Matters

Impact on research. The paper pushes evaluation away from treating benchmark scores as evidence of an underlying cognitive trait and toward task-distribution-based measurement with explicit difficulty structure. If adopted, it would change how benchmark suites are designed, how results are aggregated, and how claims about progress toward AGI are phrased.

Real-world applications:

  • Model selection in deployment. Teams choosing between models for question answering, summarization, or coding would compare performance breadth and reliability across a task distribution rather than a single headline benchmark score.
  • Benchmark design for evolving use cases. Because generality-based evaluation does not require pre-defining a canonical "core" task set, new tasks can be added as they emerge rather than waiting for a fixed suite to be revised.
  • Robustness testing under task drift. The agent-characteristic curve makes it visible whether a model fails mostly on trivial tasks or mostly on the hardest ones — two profiles that a mean score would treat as identical.
  • Multitask and instruction-tuned training. The √n reduction provides a formal rationale for the diverse-prompt, cross-task training already used in models like T0, FLAN-T5, and OPT-IML.

Industry relevance. Static benchmarks saturate quickly and can be gamed by test-specific heuristics, so companies relying on leaderboard position as a proxy for product quality face misalignment between measured and delivered capability. A generality framing ties evaluation to deployment conditions — varied, shifting, and open-ended — which is closer to how commercial systems actually encounter tasks.

Future Directions

  • Operationalizing the generality measure. The paper defines Γ_M = 1/S_M formally but does not specify how practitioners should estimate the task distribution Q or the difficulty function h(t) in concrete benchmark suites.
  • Building benchmark suites around task distributions. Moving from fixed challenge sets to sampled, difficulty-stratified task suites would require new infrastructure and reporting conventions that the paper does not detail.
  • Testing the empirical claim beyond the two figures. The correlation results are presented for ARC-AGI-1, ARC-AGI-2, LMArena, OpenBookQA, Entity Extraction, and StackUnseen; broader validation across other benchmark families and model sets is left open.
  • Reconciling generality with domain-specific systems. The paper argues narrow systems are brittle under task drift, but does not specify when a specialized system is nevertheless the right choice or how generality and specialization should be traded off.

Target Audience

Researchers and practitioners in AI evaluation and benchmarking, machine learning theorists interested in multitask generalization bounds, and philosophers of AI working on measurement and the concept of intelligence. It is also useful for product and engineering teams who select models based on benchmark scores and want a principled alternative to leaderboard-driven decisions. Readers without background in PAC learning will find the conceptual sections accessible, while the formal sections assume familiarity with statistical learning theory.

Authors’ abstract

Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are widely used to evaluate the intelligence of large language models (LLMs). Yet, the concept of intelligence remains elusive- lacking a stable definition and failing to predict performance on practical tasks such as question answering, summarization, or coding. Optimizing for such benchmarks risks misaligning evaluation with real-world utility. Our perspective is that evaluation should be grounded in generality rather than abstract notions of intelligence. We identify three assumptions that often underpin intelligence-focused evaluation: generality, stability, and realism. Through conceptual and formal analysis, we show that only generality withstands conceptual and empirical scrutiny. Intelligence is not what enables generality; generality is best understood as a multitask learning problem that directly links evaluation to measurable performance breadth and reliability. This perspective reframes how progress in AI should be assessed and proposes generality as a more stable foundation for evaluating capability across diverse and evolving tasks.

Read the original paper