Skip to content
AI.info

Research

Incoherence in Goal-Conditioned Autoregressive Models

Incoherence in Goal-Conditioned Autoregressive Models Overview Research area: Reinforcement learning theory — specifically the control-as-inference framework, goal-conditioned generative policies, and

Incoherence in Goal-Conditioned Autoregressive Models
arXiv
2510.06545
Published
2025-10-08
Authors
Jacek Karwowski, Raymond Douglas

AI summary

Incoherence in Goal-Conditioned Autoregressive Models

Overview

Research area: Reinforcement learning theory — specifically the control-as-inference framework, goal-conditioned generative policies, and iterative policy improvement (re-training a policy on its own actions).

Technical level: Advanced. The paper is almost entirely mathematical: it defines soft Q/V functions through a probabilistic (optimality-variable) formulation, proves equivalence theorems between three families of policy-improvement procedures, and gives convergence-rate results. Familiarity with MDPs, KL divergence, and KL-regularised RL is assumed.

Scope in one sentence: The paper redefines "incoherence" for goal-conditioned autoregressive policies and proves that iteratively re-training such a policy on its own trajectories removes incoherence, monotonically improves return, and coincides exactly with two other update schemes (posterior-folding into the reward, and temperature reduction) when the environment is deterministic.

Authors: Jacek Karwowski (Department of Computer Science, University of Oxford) and Raymond Douglas (Telic Research). arXiv ID 2510.06545.

What This Paper Is About

When a generative model over actions is conditioned on a goal ("reach reward R") and then used autoregressively, step by step, the resulting policy answers the wrong question. It picks the action that would succeed if future choices were made by the original prior, rather than the action that would succeed if future choices were made by the goal-conditioned policy itself. This mismatch is called incoherence, and it is structural — it is not fixed by making the underlying model more accurate.

The paper formalises this notion, shows how it relates to (and differs from) optimality, and then analyses the widely used fix of fine-tuning the model on its own rollouts. The main goal is to characterise exactly what that re-training trajectory looks like, how fast it improves, and how it connects to other, seemingly different, ways of tightening a goal-conditioned policy.

Key Contributions

  1. A refined and generalised definition of incoherence. The authors replace the earlier notion from Douglas et al. (2024) with an f-incoherence (Definition 4.6): the KL divergence over trajectories between the policy and its own f-soft Q policy, for an arbitrary "order-respecting" function f (Definition 4.4). This separates the concept of coherence from optimality — a policy can be coherent without being optimal.

  2. A three-way correspondence for removing incoherence. The paper shows that three apparently distinct procedures produce the same sequence of policies: (a) control-as-inference re-training, where the model is iteratively conditioned on its own rollouts (Definition 5.3); (b) folding the posterior over actions into the reward, in the spirit of Levine's (2018) prior-folding trick (Definitions 5.7 and 5.8); and (c) reducing the temperature / increasing the inverse temperature of the optimality variable (Definition 5.6).

  3. Equivalence and separation theorems. For deterministic dynamics and inverse temperature α(k) = 2^k, all three sequences coincide exactly (Theorem 5.9). For stochastic dynamics, only the first two coincide; a counterexample (Example 3) shows temperature-based updating diverges from the others.

  4. Rates and convergence. From the equivalence, the authors derive a return-improvement rate for re-training (Corollary 5.10) and show that incoherence vanishes in the joint limit of low temperature and many re-training steps (Corollary 5.11).

Main Findings

  • Naive goal-conditioning is not coherent. The "mountain race" example (Example 1) makes this concrete. From state ∅ the agent chooses between a risky path ↗ and a safe path ↘. With a uniform prior over actions, conditioning on success assigns π(↗|∅) ∝ 1/2 and π(↘|∅) ∝ 3/4 — because the conditioned policy evaluates the risky first move against the prior's later behaviour, not against its own. A coherent policy would know that arriving in the risky branch it will itself choose correctly.

  • Iterated coherence converges in at most T steps. The iterated f-coherence procedure π^ℬ (Definition 4.9) — effectively soft value iteration with an f-transformed soft Q — produces an f-coherent policy after at most T steps (Proposition 4.10). In the mountain-race example with δ = 1, the fixed point is reached after T = 2 iterations.

  • Coherence implies a limiting element of optimality, but only as a necessary condition. As δ → 0, the Boltzmann-rational policy converges in distribution to the uniform distribution over the set of maximal-Q actions (Proposition 4.13). If the prior is itself optimal, the limit is optimal (Corollary 4.14). But lim_{δ→0} κ_δ(π) < ∞ is only necessary for optimality, not sufficient (Corollary 4.15): a prior that assigns zero probability to ↗ in the root state can be Q-greedy everywhere except at t = 0.

  • Re-training improves return monotonically. The control-as-inference sequence satisfies J(π^𝒢_{i+1}) ≥ J(π^𝒢_i) (Proposition 5.4, "strong return improvement lemma"), and the sequence converges to a limit policy (Proposition 5.5, restating Douglas et al., 2024, Theorem 3.6). If the prior has full support over all actions in all states, that limit is optimal.

  • Three updates are the same update — deterministically. With deterministic transitions, for all k ∈ ℕ₊ and α(k) = 2^k: π_{α(k)} = π^𝒢_{2^k} = π^ℱ_{2^k} = π^ℋ_k (Theorem 5.9). Under stochastic dynamics only π^ℱ_k = π^𝒢_k holds.

  • Stochasticity breaks the temperature correspondence. In a three-state, two-action MDP with uniform prior and rewards r(s₁) = log(1/3), r(s₂) = log(2/3), the first re-training step gives π^𝒢_1(a₁|∅) = 25/61, while the temperature-updated policy gives π_{α(2)}(a₁|∅) = 7/17 (Example 3).

  • The improvement rate is explicit and temperature-governed. Corollary 5.10 expresses J(π^𝒢_k) − J(π^𝒢_{k−1}) as a ratio involving the derivative of the return and the causal entropy of the policy, divided by quantities including the second derivative, plus an O(1/k²) remainder — all derivatives taken with respect to temperature α.

  • Incoherence disappears at the limit. lim_{(δ,i)→(0,∞)} κ_δ(π^𝒢_i) = 0 (Corollary 5.11).

  • The correspondence has computational content. The paper connects it to the training–inference trade-off: reducing temperature in information-bounded RL can be implemented via top-k rejection sampling, such as Speculative Rejection (Sun et al., 2024). The text is truncated at this point.

Methodology in Plain English

The authors work with ordinary Markov decision processes — states, actions, transitions, an initial distribution, a non-positive reward function, and a discount factor (assumed to be 1 without loss of generality, since any discounted MDP can be converted by adding an auxiliary terminal state).

Their central move is to treat the reward as a probability rather than a score. A non-positive reward r(s, a) is reinterpreted as log p(𝒪 = 1 | s, a), where 𝒪 is a binary "optimality" variable. This is the standard control-as-inference device, but with one crucial modification: when computing the soft value function V, the expectation over actions is taken under the policy π itself, not under a fixed prior (Definition 4.1). That single change is what makes the definition "coherent" — it forces the value to reflect the actions the agent will actually take later.

Incoherence is then simply the KL divergence between a policy's trajectory distribution and the trajectory distribution induced by its own soft-Q-derived policy (Definition 4.6), which factorises into a sum over timesteps of per-state KL divergences weighted by the state occupancy measure (Proposition 4.7).

To study the fixes, the authors define three operators on policies: repeated goal-conditioning (control-as-inference), repeated posterior-folding into the reward, and conditioning on a reward scaled by a rising inverse temperature. They then prove these operators commute or coincide under stated conditions, which lets results proved for one formulation transfer to the others — for example, letting them read off the convergence rate of re-training from the causal-entropy characterisation of the temperature-annealed policy.

The paper also ships a code appendix implementing tabular MDP experiments at github.com/jkarwowski/incoherence, which numerically validates the main results. No large-scale empirical benchmark results are reported in the paper text.

Why This Matters

Impact on research. The paper gives a structural, not statistical, account of a failure mode that recurs across return-conditioned and goal-conditioned methods. It connects the theory of Decision Transformer-style architectures and return-conditioned supervised learning (Brandfonbrener et al., 2023; Srivastava et al., 2021; Chen et al., 2021), where trajectory "luck" in stochastic environments causes systematic failure (Štrupl et al., 2022; Paster et al., 2022; Yang et al., 2024), to KL-regularised policy search (Peters et al., 2010; Schulman et al., 2017; Abdolmaleki et al., 2018; Korbak et al., 2022) and to residual energy-based text generation (Deng et al., 2020). It also latches incoherence to the notion of effective horizon from Laidlaw et al. (2024). Notably, the paper distinguishes its use of "incoherence" from O'Donoghue et al. (2020), who used the term for a different problem about epistemic uncertainty in the exploration–exploitation trade-off.

Real-world applications (as motivated by the paper's framing):

  • Fine-tuning language models into agents. The paper reasons about LLMs that simulate agents (Shanahan et al., 2023; Douglas et al., 2024), including the argument from Andreas (2022) that any model trained on human internet behaviour is effectively goal-conditioned on human outcomes. The incoherence problem applies directly to prompting and scaffolding methods that condition a base model purely formally.

  • RLHF and preference-based fine-tuning pipelines. The paper flags the connection to RLHF (Christiano et al., 2017), DPO (Rafailov et al., 2023) and GRPO (Shao et al., 2024), and cites the "RLHF Conditioning Hypothesis" (Hubinger et al., 2023) as an open question this framing bears on.

  • Inference-time search and rejection sampling. Because temperature reduction in the paper's framework maps onto best-of-n rejection sampling and top-k methods like Speculative Rejection (Sun et al., 2024), the equivalence results speak to how much of the benefit of search can be absorbed into the model rather than paid for at inference time.

  • Game-playing and expert iteration systems. Expert iteration and MCTS underpin AlphaZero (Silver et al., 2017; Silver et al., 2018) and MuZero (Schrittwieser et al., 2020), with the general formulation by Anthony et al. (2017). The paper notes prior work did not address the soft-conditioning case that this analysis covers.

Industry relevance. If a returned-reward-conditioned or goal-conditioned policy is deployed autoregressively, the paper says, the failure is not a data problem or a scale problem — it is an objective mismatch that only re-training on the model's own actions (or its equivalents) repairs. That is directly actionable for teams building return-conditioned decision models and for those deciding whether to spend compute on training-time iteration or inference-time sampling.

Future Directions

  • What the stochastic case loses. Theorem 5.9 shows temperature reduction and re-training diverge under stochastic dynamics, and Example 3 is the pointed demonstration. Characterising when annealing can substitute for re-training — and bounding the gap when it cannot — is a natural next step. Example 3's construction is offered as a template.

  • The effective-horizon link, developed further. The abstract states that incoherence connects to the effective horizon of Laidlaw et al. (2024) through soft-conditioning generative models; the full development sits in the (truncated) later sections and Section 6, which also covers limitations.

  • Turning the necessary condition into something usable. Corollary 4.15 gives only a necessary condition for optimality via lim_{δ→0} κ_δ(π) < ∞. Whether a practical sufficient condition exists — and what it would require of the prior — is left open.

  • From tabular validation to language models. The paper's setup considers a model trained on a single environment, and it explicitly notes that the exact nature of formal conditioning in LLMs, and its relationship to RL fine-tuning, remains an open problem, alongside the RLHF Conditioning Hypothesis. Extending the equivalence results beyond tabular MDPs to the LM setting is the obvious frontier.

Target Audience

Researchers in reinforcement learning theory and probabilistic machine learning, particularly those working on control-as-inference, KL-regularised policy optimisation, soft Q-learning, and soft actor-critic (Haarnoja et al., 2017; Haarnoja et al., 2018). Also relevant to practitioners building return-conditioned or goal-conditioned policies (Decision Transformer, RvS, upside-down RL) and to those studying the theory of how language models are steered toward goals via conditioning and RL fine-tuning. Readers need comfort with MDP formalism, KL divergences, and variational/energy-based derivations; the paper is not an introduction to these topics.

Authors’ abstract

We investigate mathematically the notion of incoherence: a structural issue with reinforcement learning policies derived by naive goal-conditioning of autoregressive models. We focus on the process of re-training models on their own actions, that is, fine-tuning offline-learned policies with online RL. We prove that it decreases incoherence and leads to an improvement in return, and we aim to characterize the resulting trajectory of policies. By re-framing standard notions of control-as-inference and soft Q learning, we establish a three-way correspondence with two other ways of understanding the iterative re-training process: as folding the posterior into the reward and, in the deterministic case, as decreasing the temperature parameter; the correspondence has computational content via the training-inference trade-off. Through soft-conditioning generative models, we discuss the link between incoherence and the effective horizon.

Read the original paper