Skip to content
AI.info

Research

Projected Autoregression: Autoregressive Language Generation in Continuous State Space

Overview Research area: Natural language processing, specifically autoregressive language generation and decoding algorithms for Transformer language models. Technical level: Intermediate. The paper a

Projected Autoregression: Autoregressive Language Generation in Continuous State Space
arXiv
2601.04854
Published
2026-01-08
Authors
Oshri Naparstek

AI summary

Overview

Research area: Natural language processing, specifically autoregressive language generation and decoding algorithms for Transformer language models.

Technical level: Intermediate. The paper assumes familiarity with autoregressive decoding, embeddings, and contrastive training objectives, but the core idea is intuitive and the empirical setup is small-scale and controlled.

Scope: A single-author study that proposes and evaluates "Projected Autoregression," an autoregressive interface that predicts continuous token vectors in embedding space and defers discrete token commitment, tested via LoRA adapters on a Granite 3B Code Base backbone trained on FineWeb.

What This Paper Is About

Standard autoregressive language models must pick a discrete token at every step, which fuses prediction with an irreversible commitment. This paper asks whether that coupling is necessary, and shows that generation can instead predict a continuous vector and only convert it into a token when commit time arrives. The goal is to characterize what this change does to the text a model produces, and what new control knobs it exposes.

Key Contributions

  1. A new autoregressive interface. Projected Autoregression replaces next-token selection with continuous prediction in embedding space, followed by discrete projection (nearest-neighbor commit) only at commitment time. Discrete tokens arise as a downstream interface rather than as the prediction target.

  2. Delayed commitment via a "liquid tail." An optional short mutable suffix of length K lets the model iteratively refine tail vectors before they are committed, keeping refinement local and causal rather than a sequence-wide denoising process.

  3. A demonstrated regime shift localized to continuous prediction. A K=1 decomposition (immediate projection, no refinement) shows the change in text structure already appears before multi-step refinement, and a compute-matched best-of-16 autoregressive baseline does not reproduce it.

  4. A continuous control surface. The paper identifies five knobs — direction rate, history noise, tail length K, state-space classifier-free guidance, and embedding geometry — that act on the evolving generative state rather than on token probabilities.

Main Findings

  • Projected Autoregression occupies a different operating regime. On the shared Granite 3B backbone and FineWeb data (50 prompts, 150 tokens/steps), Projected AR at K=16 with deterministic argmax reaches srep4 .001, d-1 .911, MAUVE .849, and info-d 12.0. AR greedy gets .480 / .314 / .218 / 6.9; AR greedy+rep gets .000 / .873 / .953 / 10.1; AR top-p=0.95 gets .006 / .693 / .946 / 9.2; compute-matched AR best-of-16 gets .058 / .624 / — / 9.0. Projected AR leads on repetition, diversity, and information density, while AR heuristics lead on MAUVE.

  • The regime shift is present at K=1. With immediate projection (no refinement), srep4 is .002, d-1 .888, info-d 13.8, discourse markers 0.68, connectives 2.02, question words 2.67. Moving to K=16 changes these to .001 / .911 / 12.0 / 0.54 / 1.84 / 2.57, adding diversity and further reducing repetition — so the shift is attributable to the continuous prediction interface, not only to delayed commitment.

  • Neighborhood narrowing is the mechanism. As tail vectors mature, SpreadNet (cosine distance to the centroid of top-k nearest vocabulary embeddings) decreases monotonically and CentroidNet (cosine similarity to that centroid) increases. From K=2 to K=16 the mean probe change goes from −0.037 to −0.088 (SpreadNet) and +0.080 to +0.181 (CentroidNet), with the appendix K-sweep indicating saturation around K≈16. The tail converges toward the centroid of a narrowing cloud of plausible tokens, not toward a single embedding.

  • Embedding geometry controls register stability. With frozen pretrained embeddings, generated text drifts mid-generation from expository to promotional or organizational voice. In a manual annotation of 10 prompts per condition, learned embeddings raise register-consistent outputs from 4/10 to 8/10 at K=1 and from 2/10 to 8/10 at K=16, and reduce register drift from 3/10 and 5/10 to 0/10 in both cases. The effect persists at K=1, so drift is a property of token-space geometry rather than of delayed commitment. The cost is slightly less discourse-rich text, and roughly a 30% reduction in attribution-like patterns.

  • Direction rate has a sweet spot. Sweeping (η_min, η_max) from (0.01, 0.10) to (1.0, 1.0) moves discourse markers from 0.22 to 0.70 per 100 words at the optimal setting (0.10, 0.75), a 3× range. Too-slow rates produce short, incoherent output (46 words) with artificially inflated diversity (d-1 = .986) and info-d 13.2 flagged as an artifact; too-fast rates eliminate refinement (disc. = 0.20).

  • History noise gives sampling-free diversity. At σ_h = 0.5, content changes substantially across runs while srep4 stays at .000 and d-1 at .910, and discourse structure is preserved. Beyond σ_h = 0.5 the prefix is too corrupted for coherence.

  • Text quality is flat across K. For K ∈ {2, 4, 8, 16, 32}, all configurations show srep4 ≤ .002 and d-1 ≥ .88. Warmup matters mainly at K=32, where output length rises from 75 to 108 words.

  • Overhead is optional. The liquid tail adds K-fold per-token overhead (recomputing K positions per step while reusing the prefix KV cache). Because K=1 already produces most structural benefits, practitioners can use K=1 for no overhead.

Methodology in Plain English

The authors keep a standard causal Transformer backbone but change what it predicts. Instead of outputting scores over a fixed vocabulary, the model regresses a continuous vector in a fixed embedding space, using an MSE loss alongside an InfoNCE contrastive loss (final objective: 0.2 × MSE + 1.0 × NCE, weighted by 1−α_t+0.1) to keep vectors anchored to discrete token identities and prevent collapse toward frequent tokens.

Each position carries a maturity value α between 0 and 1 — near 0 for freshly introduced tokens, near 1 for tokens close to commitment. The model is conditioned on α through a sinusoidal encoding plus MLP, and on the tail length K through FiLM modulation. During training, the model is shown artificially corrupted tail vectors (α_t · true embedding + (1−α_t) · noise, re-projected onto a fixed-radius sphere), with the tail length sampled uniformly from 1 to K_max, 20% of history tokens randomly dropped to α=0, tail length occasionally replaced by an "unknown" token with 25% probability, and optional α jittering. Two attention modes are alternated 50/50: a full causal mask, and a tail-only mask where tail positions cannot see the prefix — the latter trains the unconditional branch used for classifier-free guidance.

At inference, generation runs left to right. A liquid tail of length K is initialized with random low-norm vectors. Each step produces a context-aware and a context-masked prediction, combined as ẑ = ẑ_cm + s(ẑ_ca − ẑ_cm) with guidance scale s. Tail vectors are updated by a contraction step z̃_i ← z̃_i + η(α_i)(ẑ_i − z̃_i), where η increases with maturity. When a vector reaches the front of the tail, it is committed by argmax over cosine similarity to the embedding matrix and replaced by that embedding. Optimization uses AdamW with separate learning rates for LoRA and optional embedding parameters, batch size 1 with 16 gradient accumulation steps, gradients clipped to norm 1.0, and 50K–600K training steps depending on embedding configuration.

The evaluation applies LoRA (rank 8, α=32) to Granite 3B Code Base, trained on FineWeb, and compares against a standard AR LoRA adapter with identical configuration and data across 50 prompts and 150 tokens/steps per response.

Why This Matters

Impact on research. The paper reframes token selection as one point in a larger family of autoregressive interfaces, alongside continuous-latent approaches like Coconut, hybrid drafting like TiDAR, chunk-level continuous prediction like CALM, and global-denoising diffusion language models. It argues for a distinction between local causal refinement and sequence-wide denoising, and shows that a continuous-state interface is compatible with strict left-to-right causality.

Real-world applications (as suggested by the work, not demonstrated at deployment scale):

  • Diversity-controlled generation — history noise produces different content across runs while preserving text structure metrics, offering a sampling-free alternative to token-level sampling.
  • Longer or more varied open-ended writing — the anti-degeneration behavior at K=16 (srep4 .001) addresses repetition, a common failure mode in deterministic decoding.
  • Style- and register-stable assistants — learned embedding geometry reduced register drift to 0/10 in the reported manual annotation, relevant where an assistant must not shift voice mid-response.
  • Steerable generation — direction rate and state-space guidance offer controls that act on the generative state rather than post-hoc token probabilities.

Industry relevance. For teams already serving Transformer models, the practical question is overhead versus benefit: K=1 requires no extra per-token compute and already produces the reported structural differences, while K>1 buys extra diversity at K-fold per-token cost. The paper notes prompt relevance is lower than AR, which is an important caveat for deployment.

Future Directions

  • Moving beyond the 3B LoRA setting. The main study is controlled — same backbone, data, and adaptation budget — but evaluated at 3B LoRA scale on 50 open-ended prompts. Whether the regime shift holds at larger scale or full fine-tuning is not established, though appendix GPT-2 experiments suggest the phenomenon extends beyond Granite.
  • Closing the prompt-relevance gap. The paper reports that prompt relevance is lower than AR; recovering it while keeping the diversity and anti-degeneration benefits is an open problem.
  • Disentangling the richness–stability trade-off. Learned embeddings eliminate register drift but yield slightly less discourse-rich text. The paper does not resolve whether this trade-off can be avoided.
  • Understanding saturation and scheduling. Neighborhood narrowing saturates around K≈16, and the growing-tail effect (the tail starts shorter than K and grows) dilutes large K values. Better schedules, warmup strategies, and conditioning schemes are natural next steps.

Target Audience

Researchers and engineers working on language model decoding, generation algorithms, and continuous or latent-space language modeling. It is most useful to readers already comfortable with autoregressive Transformers, embedding spaces, and contrastive objectives who want to understand what changes when prediction and commitment are separated. Practitioners seeking a drop-in decoding improvement will find the K=1 no-overhead option and the control-knob survey most actionable, while those evaluating the approach for production should weigh the lower prompt relevance and the single-backbone, 50-prompt evaluation scope.

Authors’ abstract

Standard autoregressive language models generate text by repeatedly selecting a discrete next token, coupling prediction with irreversible commitment at every step. We show that token selection is not the only viable autoregressive interface. \textbf{Projected Autoregression} replaces token selection with continuous prediction in embedding space followed by discrete projection at commitment time. The model predicts next-token vectors via regression and contrastive objectives, while discrete tokens arise only by nearest-neighbor projection. An optional mutable suffix (``liquid tail'') enables iterative refinement before commitment, but the central change is more basic: next-step prediction is continuous, and discrete tokens are produced only as a downstream interface. Projected Autoregression establishes a concrete alternative to token-selection autoregression: language generation can be organized around continuous-state prediction with delayed discrete commitment. Refinement remains local to a short causal suffix within a left-to-right causal process, rather than a sequence-wide denoising process. This separation has two consequences. First, it induces a \emph{distinct generation regime}: even with immediate projection ($K{=}1$), continuous prediction yields text structure and dynamics that differ from tested token-space AR baselines, including a compute-matched best-of-16 reranking baseline. Second, it exposes a \emph{continuous control surface} inside autoregressive generation: direction rate, history noise, delayed commitment, state-space guidance, and embedding geometry act directly on the evolving generative state before token commitment. Taken together, these results place repeated token selection within a larger family of autoregressive interfaces and expose continuous state space as a broader algorithmic design space for language generation.

Read the original paper