Skip to content
AI.info

Research

Do LLMs Understand Sequential Structure? A Controlled Study of Inference and Generation

Overview Research area: Evaluation of large language models as sequential reasoners and behavioral simulators, combining controlled game-based probing (Rock–Paper–Scissors) with stochastic n-gram cont

Do LLMs Understand Sequential Structure? A Controlled Study of Inference and Generation
arXiv
2610.04977
Published
2026-10-04
Authors
Jerry Wang, Zhengxiang Wang, Ting Yu Liu, Hsin-Ling Hsu, Yi-Cheng Lai, Tengfei Ma

AI summary

Overview

  • Research area: Evaluation of large language models as sequential reasoners and behavioral simulators, combining controlled game-based probing (Rock–Paper–Scissors) with stochastic n-gram continuation tasks.
  • Technical level: Intermediate. Readers should be comfortable with basic notions of conditional probability, Markov rules, in-context learning, and next-token generation, but the paper is built around plain, small action spaces and readable metrics.
  • Scope (one sentence): The paper asks whether LLMs can recover and then faithfully reproduce the latent conditional process behind an observed sequence, rather than merely matching its surface action frequencies.

What This Paper Is About

Observed behavior does not uniquely determine what generated it: two players can produce nearly identical action proportions while following completely different conditional rules (one sampling independently, the other reacting to recent history). The authors build controlled environments where the true generating process is known exactly, so they can separate three abilities that are usually conflated in agent evaluations: identifying latent strategies, matching marginal action distributions, and executing conditional rules during generation. The goal is to test whether apparent behavioral fidelity actually reflects recovery of sequential structure.

Key Contributions

  1. A controlled two-player Rock–Paper–Scissors framework with a closed, fully specified candidate pool (16 non-Markovian distribution players labeled A–P, plus first-order opponent-only players X/Y/Z, first-order joint-state players Q/R/S, second-order joint-state players T/U/V, and second-order opponent-only players x/y/z) that lets the authors score strategy identification, distributional fidelity, and rule execution separately.
  2. A progression of three RPS experiments moving from strategy identification (Experiment 1, seven model configurations, context lengths T ∈ {100, 200, 500, 1000}) to first-order Markov inference-and-simulation (Experiment 2, four core configurations, 1000 generated rounds) to joint-state and second-order structures (Experiment 3, using deepseek-reasoner).
  3. A one-player stochastic n-gram continuation task over the abstract vocabulary {A, B, C} with dependency orders n ∈ {1, 2, 4, 8}, which removes player attribution and Red–Paper-style semantics, and compares prefix-only prompting against rule-given prompting.
  4. New evaluation metrics for stochastic high-order generation — context-conditioned likelihood gain (CCLG) and state-frequency-weighted Jensen–Shannon divergence (WJS) — alongside exact-match accuracy, rule overlap, strict rule match, cumulative strict match, MSE, and the relative Markov–Non-Markov MSE gap of Eq. 3.

Main Findings

  • Markov strategies are harder to identify than non-Markov strategies. Identification accuracy is lower on average for Markov players across all seven model configurations, and the gap is significant in six of seven cases (p < .001). GPT-5 is the sole exception (p = .095).
  • More context does not help identification. Increasing the observed history from 100 to 1000 rounds fails to improve accuracy and often degrades it: all seven models decline on non-Markov players and nearly all decline on Markov players. A non-LLM maximum-likelihood baseline reaches 100% accuracy at every context length under the same closed candidate pool, so the failures are not due to missing evidence. Appendix G also reports that longer input length predicts lower accuracy, while output length does not reliably improve performance.
  • Correct identification improves rule-following but is not equivalent to it. Models simulate the target rule better when the identity is correct, especially under strict rule match. Yet wrong-identity cases do not collapse to chance — rule overlap stays above the random baseline in several models — and cumulative strict rule match sometimes meets or exceeds identity accuracy.
  • Models simulate statistical players better than Markov players. The Markov–non-Markov MSE gap is about 8–9% overall, but exceeds 100% in the correct-identity subset. In the incorrect-identity subset the pattern reverses, with wrong Markov predictions producing lower MSE than wrong non-Markov predictions.
  • Surface fidelity can hide a wrong mechanism. Among 282 incorrectly recovered strategies, 41.8% remain within TV ≤ 0.02 of the target marginal distribution, and 62.1% remain within TV ≤ 0.10.
  • Rule following is largely determined early. Cumulative strict rule match curves show stronger models declining mildly in early rounds and then stabilizing, while weaker models stay mostly flat, rather than degrading progressively over long generations.
  • Teacher forcing separates two failure sources. Under ground-truth histories, DeepSeek Reasoner reached 93.3% (97.2% on the correct-ID subset) and GPT-5 reached 82.5% (89.5%), implying their free-running errors partly come from sustaining the rule over time. DeepSeek Chat (52.5%, 66.7%) and GPT-5-mini (54.2%, 70.0%) still struggle with ground-truth histories, so local rule application itself remains a bottleneck for them.
  • Generation does not systematically change recognition. Comparing Experiment 1 (identification only) with Experiment 2 (identification plus generation) shows no systematic effect; no difference is significant under two-proportion z-tests.
  • Longer history, not joint-state conditioning, drives degradation. Adding joint-state information does not reduce rule-following performance, whereas both second-order settings show lower strict rule match than first-order settings even when rule overlap stays relatively high. Identity accuracy remains high in these settings, so the harder part is execution, not recognition.
  • High-order degradation persists without player attribution. In the one-player task, CCLG is often positive for n = 1 to 4, but at n = 8 it drops sharply for most models while WJS increases. A prefix-only n-gram MLE baseline remains stable at higher orders, confirming the information is recoverable.
  • Context use and rule recovery are distinct. Higher CCLG does not always correspond to lower WJS; models can show strong context-conditioned gains while still diverging from the true transition distributions.
  • Neither longer prefixes nor explicit rules fix it. Increasing prefix length from 256 to 2048, and providing the n-gram transition rules, do not consistently remove high-order degradation.

Methodology in Plain English

The authors start with Rock–Paper–Scissors because it is a tiny, fully enumerable action space where every "player" is a known, published rule. Some players are simple frequency players — always Rock, always Scissors, uniform random, or one of a set of biased distributions. Others are reactive: they look at the opponent's previous move (Win-Last, Lose-Last, Copy-Last), at both players' previous moves (finding the action neither played last round, and beating or losing to it), or at the last two rounds. The model sees a trajectory of moves and must name each player's identity, optionally estimate action distributions, and then continue the game for many rounds without being told who is who.

Because two players can look statistically similar but behave differently, the authors score the continuation on two different axes. For reactive players they check rule overlap — whether each generated move equals what the true rule prescribes given the generated history — plus strict rule match within windows and cumulative strict match over time. For frequency players they compare distributions via MSE, and they report the percentage gap between Markov and non-Markov error. They also run a teacher-forced diagnostic where the model sees the real history and predicts only the next move, which isolates immediate rule application from long-horizon drift.

The final experiment strips out Rock–Paper–Scissors and the opponent entirely. A single-player sequence is generated from a stochastic n-gram process over {A, B, C} with dependency orders 1, 2, 4, and 8. Because the process is probabilistic, exact match is meaningless, so the authors use CCLG (improvement over a unigram baseline) and WJS (state-frequency-weighted divergence from the true conditional distribution), and they compare a prefix-only setting against a rule-given setting.

Why This Matters

Impact on research. The paper argues that benchmark success in agent and simulation settings can be misleading: matching aggregate behavior, or even naming the right strategy, does not establish that a model has recovered the conditional mechanism. It supplies a diagnostic decomposition — recognition, marginal matching, conditional execution — that can be reused to re-evaluate claims about LLM social intelligence, opponent modeling, and in-context rule learning, and it shows the failure is not attributable to insufficient information, since a maximum-likelihood baseline solves the same identification tasks perfectly.

Real-world applications:

  • Evaluating LLM-driven user simulators, where a simulated user must respond to what actually happened earlier in a session rather than replaying a plausible-looking distribution of clicks or replies.
  • Multi-turn tool-use and agent pipelines, where an action two or three steps back determines the correct next call and where drifting off the intended control rule silently breaks downstream tasks.
  • Synthetic data generation and scenario testing, where a simulator that only matches marginal statistics produces convincing but structurally wrong trajectories.
  • Behavioral modeling and human-subject simulation, where the goal is to reproduce history-dependent decision patterns rather than average response rates.

Industry relevance. Teams that ship conversational agents, recommendation-style interactions, or simulated-user test harnesses rely on the assumption that a model which looks right on aggregate metrics is following the right process. This work shows that a model can sit within tight total-variation distance of the target distribution (41.8% of wrong recoveries within TV ≤ 0.02) while implementing the wrong mechanism, which is exactly the kind of failure that passes offline dashboards and only surfaces in long interactions.

Future Directions

  • Extending the framework beyond the predefined candidate pool Π to larger or learned strategy spaces, so models must infer behavioral structure without a closed set of labeled options (the paper's stated limitation).
  • Determining why second-order rules produce the largest degradation: whether the problem is composing local rule fragments, tracking longer state, or something about how transition structure is maintained during decoding.
  • Closing the gap between context use and rule recovery, since higher CCLG did not imply lower WJS; the paper's appendices R–U only begin to separate likelihood effects, output artifacts, and rule-access effects.
  • Testing whether targeted interventions — alternative trajectory representations, output schemas, or decoding settings — can improve sustained execution of high-order conditional structure, given that prompt variants and explicit rules did not resolve it.

Target Audience

Researchers and practitioners evaluating LLMs as agents, simulators, or sequential reasoners; people working on in-context learning, opponent modeling, and behavioral simulation; and engineers building multi-turn systems who need evaluation methods that distinguish distribution matching from genuine conditional rule following. Readers looking for a general introduction to LLMs will find the setting accessible, but the metrics, baselines, and appendix-level diagnostics assume some familiarity with probabilistic sequence modeling.

Authors’ abstract

Large language models (LLMs) are increasingly used as interactive agents and simulators, yet it remains unclear whether they can recover latent sequential structure beyond surface action frequencies. This distinction is critical for behavioral simulation, where actions are often shaped by prior context rather than marginal frequencies alone. We study this question using controlled two-player Rock--Paper--Scissors interactions and a one-player stochastic n-gram continuation task. Across these experiments, we test whether LLMs can identify latent strategies, follow simple Markov rules, and sustain higher-order conditional dependencies. Our framework separates distribution matching from conditional rule following. Results show that longer context does not improve identification, correct recognition does not ensure faithful simulation, and higher-order dependencies substantially degrade rule recovery. Apparent behavioral fidelity can therefore mask incorrect generative mechanisms.

Read the original paper