Skip to content
AI.info

Research

The Idola Tribus of AI: Large Language Models tend to perceive order where none exists

Overview Research area: Natural Language Processing — evaluation of logical reasoning, self-consistency, and cognitive bias in large language models (LLMs). Technical level: Intermediate. The experime

The Idola Tribus of AI: Large Language Models tend to perceive order where none exists
arXiv
2510.09709
Published
2025-10-10
Authors
Shin-nosuke Ishikawa, Masato Todo, Taiki Ogihara, Hirotsugu Ohba

AI summary

Overview

  • Research area: Natural Language Processing — evaluation of logical reasoning, self-consistency, and cognitive bias in large language models (LLMs).
  • Technical level: Intermediate. The experimental design is simple to follow (number series plus prompts), but it assumes familiarity with chain-of-thought reasoning, "thinking" models, and LLM-as-a-Judge evaluation.
  • Scope in one sentence: The paper tests whether five high-performance LLMs, including multi-step "thinking" models, invent patterns that contradict the numbers they are given when asked to explain integer sequences.

What This Paper Is About

LLMs are increasingly trusted to reason step by step, yet their usefulness depends on whether they stay logically consistent with the information they receive. The authors ask LLMs to describe the regularity behind 724 integer series, ranging from simple arithmetic sequences to random numbers with no order, and check whether the explanations actually fit the numbers. They find that models frequently claim to see a pattern even in purely random series, a tendency the authors call the LLM equivalent of Idola Tribus — Francis Bacon's term, from Novum Organum (1620), for the human habit of supposing more order in the world than exists.

Key Contributions

  1. A simple, knowledge-independent test of inductive bias. By using number series instead of knowledge questions, the experiment measures pattern recognition and abstraction separately from hallucination caused by gaps in training data.
  2. First systematic demonstration of over-recognized patterns in LLMs. The authors report false regularities generated for randomly generated series across all five models tested, which they frame as evidence of an AI analogue of Idola Tribus.
  3. Evidence that the bias survives multi-step reasoning. The tendency appears not only in standard models but also in thinking models with built-in self-evaluation — OpenAI o3, o4-mini, and Google Gemini 2.5 Flash Preview Thinking — even though these models score higher overall.
  4. A prompt-level countermeasure. Adding a prompt that explicitly permits the model to declare a series random substantially increases both random declarations and overall success rates for four of the five models, while Gemini 2.5 Flash Preview Thinking shows no comparable change.

Main Findings

  • Near-perfect performance on well-defined series. All five models scored 100% on arithmetic series (average 100%), and geometric series averaged 98.8%. Difference series averaged 85.4%, though GPT-4.1 dropped to 36.0% while o3, o4-mini, and Gemini 2.5 Flash Preview Thinking reached 99.0%, 99.0%, and 100%.
  • Sharp collapse on disrupted and random series. With the original prompt, success rates averaged 41.7% (quasi-arithmetic), 34.1% (quasi-geometric), 41.8% (quasi-difference), 35.4% (random-increasing), and 26.8% (random). Individual random-series scores ranged from 8.0% (GPT-4.1) and 0.0% (Llama 3.3) up to 52.0% (o3).
  • Thinking models outperform non-thinking models. Overall success was 73.1% (o3), 70.6% (Gemini 2.5 Flash Preview Thinking), 67.8% (o4-mini), 39.9% (Llama 3.3), and 33.0% (GPT-4.1), averaging 56.9%.
  • Models rarely said "random" unless told they could. Under the original prompt, the rate of random-series declarations was 3.8% overall (1.3% for GPT-4.1, 8.0% for o3, 4.3% for o4-mini, 5.4% for Gemini, 0.0% for Llama 3.3). Under the random-allowing prompt, this rose to 32.0% overall, reaching 75.8% on the random category alone.
  • Permitting "random" improves accuracy. Overall success with the random-allowing prompt rose to 71.3%, with o4-mini at 88.3% and o3 at 87.7%. Gemini 2.5 Flash Preview Thinking improved only from 70.6% to 68.9%.
  • o3 contradicted its own judgment. When o3 was used both to identify patterns and to judge them, it rated its own descriptions as valid in only 52 out of 100 random-category cases — meaning it produced patterns that it itself assessed as incorrect.
  • Models preferred "some pattern" over "no pattern" on disrupted series. For the best-performing models under the random-allowing prompt, success on random series was actually higher than on quasi-ordered series, where a single deviating term makes a pattern look plausible.
  • Creative but unverifiable explanations appeared. o3 and o4-mini produced interpretations based on atomic numbers, football players, piano, tarot, the Holy Bible, and telephone country codes, all inconsistent with the given numbers.
  • The bias was shared across all five models. Because Idola Tribus describes a tendency common to a group rather than an individual, the authors note this uniformity across models is itself consistent with the concept.

Methodology in Plain English

The researchers built eight categories of integer sequences: three with clean rules (arithmetic, geometric, difference), three "quasi" versions where a single term is off by +1 or −1, and two with no rule (random-increasing and random). In total there were 724 series: 81 arithmetic, 81 geometric, 100 difference, 81 quasi-arithmetic, 81 quasi-geometric, 100 quasi-difference, 100 random-increasing, and 100 random. Arithmetic series used common differences from 1 to 9 and first terms from 1 to 9; geometric series used common ratios from −5 to 5 excluding 0 and 1, with first terms from 1 to 9; random-increasing series were built by adding random integers between 1 and 10 starting from a first term between 2 and 18; random series used integers between 1 and 99. Values were generally positive integers up to 100, except in the geometric category.

Each of five LLMs — GPT-4.1, o3, o4-mini, Google Gemini 2.5 Flash Preview Thinking, and Llama 3.3 — received the same prompt showing five values from a series and was asked to describe the pattern in a single sentence. No system prompts were used. This produced 724 × 5 = 3,620 descriptions.

To evaluate those descriptions at scale, the authors used an LLM as judge rather than human annotators. The judge sorted each description into four options: a correct explanation matching the preset category, a correct explanation not matching it, an incorrect explanation, or a statement that the series is random. Options 1 and 2 both counted as success, and option 4 also counted as success for categories without a describable rule. A preliminary test on four held-out series and 21 author-annotated descriptions found o3 the most accurate judge at 81.0%, against 71.4% for Gemini 2.5 and 66.7% each for GPT-4.1, o4-mini, and Llama 3.3; o3 was therefore used as the judge. A second prompt variant explicitly allowed the model to answer that a series was random, and the whole comparison was repeated.

Why This Matters

Impact on research. The paper separates inductive bias from knowledge-based hallucination and argues that self-consistency cannot be assumed just because a model uses chain-of-thought or multi-step self-evaluation. It shows a failure mode that persists in thinking models, which the authors say suggests either insufficient verification capability or a willingness to assert regularities the model already recognizes as wrong — a point they connect to reported differences between CoT processes and human thinking.

The authors also stress the stakes for current engineering practice: retrieval-augmented generation and AI agent frameworks both depend on the logical consistency and self-coherence of the underlying model, so a bias toward forced pattern explanation can propagate into downstream task execution.

Real-world applications where this matters:

  • AI agent frameworks that plan multi-step actions rely on the model abstracting the right pattern from incomplete instructions; a false pattern can send an entire plan off course.
  • Retrieval-augmented generation and knowledge workflows, where the model must decide whether retrieved evidence actually supports a claim.
  • Business and data analysis, where a model asked to explain a trend may report a spurious regularity rather than stating the data shows none.
  • Conversational assistants used for hypothesis formation, where users ask open-ended questions and may not notice that the supplied explanation does not fit their own inputs.

Industry relevance. Any product that delegates open-ended interpretation to an LLM — analytics copilots, research assistants, decision-support tools — inherits this bias. The finding that a simple prompt change (explicitly allowing "this looks random") raised average success from 56.9% to 71.3% is directly actionable for prompt design, and the authors note the trade-off that such prompts may also reduce willingness to engage with genuinely complex reasoning.

Future Directions

  • Testing beyond five models. The authors state explicitly that their results do not guarantee the same tendency exists in all current and future LLMs, so broader model coverage is needed.
  • Better prompts and mitigation strategies. They note that only two prompt variations were tested and that more optimized prompts may yield improved results, without claiming the bias is unavoidable.
  • Fine-tuning approaches. The paper suggests that fine-tuning strategies previously proposed for improving logical reasoning, and attention to training data quality rather than quantity, could help address inductive bias, though those configurations were designed mainly for deductive reasoning.
  • Extending hallucination-mitigation ideas to logic. Mechanisms that let models say they do not know, or explain why they cannot answer — proposed for knowledge-based tasks — need investigation for logical reasoning tasks such as these.
  • Understanding the mechanism. Whether the bias stems from an implicit compulsion to always produce an answer, or from pressure toward efficient information processing during training or instruction tuning, remains an open question.
  • Applying the "cannot determine" option to quasi-ordered cases. The authors note it is unresolved how to make models recognize the absence of a rule when a single deviating value makes a pattern appear plausible.

Target Audience

This paper suits LLM evaluation researchers and practitioners who design agent frameworks, retrieval-augmented systems, or any pipeline that depends on a model's reasoning being consistent with its inputs. It is also relevant to prompt engineers looking for low-cost mitigations, and to readers interested in cognitive-bias analogies in AI. The methodology is accessible to those without deep mathematics background, though the authors caution that the absolute success-rate numbers come from an LLM judge rather than human annotation and may not be precise in absolute terms, even if the overall trend holds.

Authors’ abstract

We present a tendency of large language models (LLMs) to generate absurd patterns despite their clear inappropriateness in a simple task of identifying regularities in number series. Several approaches have been proposed to apply LLMs to complex real-world tasks, such as providing knowledge through retrieval-augmented generation and executing multi-step tasks using AI agent frameworks. However, these approaches rely on the logical consistency and self-coherence of LLMs, making it crucial to evaluate these aspects and consider potential countermeasures. To identify cases where LLMs fail to maintain logical consistency, we conducted an experiment in which LLMs were asked to explain the patterns in various integer sequences, ranging from arithmetic sequences to randomly generated integer series. While the models successfully identified correct patterns in arithmetic and geometric sequences, they frequently over-recognized patterns that were inconsistent with the given numbers when analyzing randomly generated series. This issue was observed even in multi-step reasoning models, including OpenAI o3, o4-mini, and Google Gemini 2.5 Flash Preview Thinking. This tendency to perceive non-existent patterns can be interpreted as the AI model equivalent of Idola Tribus and highlights potential limitations in their capability for applied tasks requiring logical reasoning, even when employing chain-of-thought reasoning mechanisms.

Read the original paper