Skip to content
AI.info

Research

Principled Thoughts for Latent Recursive LLM Systems

Overview Research area: Training objectives for large language models that reason in continuous latent space, including single-model latent recursion and multi-agent latent communication. Technical le

Principled Thoughts for Latent Recursive LLM Systems
arXiv
2609.36159
Published
2026-09-28
Authors
Fahd Seddik, Fatemeh Fard

AI summary

Overview

Research area: Training objectives for large language models that reason in continuous latent space, including single-model latent recursion and multi-agent latent communication.

Technical level: Advanced. The paper builds on a formal framework of four properties of a valid thought representation (Seddiq and Fard, 2026) and derives differentiable loss terms from them, so familiarity with cross-entropy training, KL divergence, and latent reasoning systems helps.

One-sentence scope: The paper proposes REST (REpresentation-Supervised Thought(s)), a training objective that adds four property-based loss terms to the cross-entropy loss of latent recursive LLM systems, and evaluates it on seven benchmarks across single-agent and multi-agent settings, model sizes, and three training seeds.

What This Paper Is About

Latent recursive LLM systems reason in continuous space rather than decoded text: a model can recur on its own hidden states, or pass those states to other agents. Training such systems typically supervises only the cross-entropy (CE) of the final decoded answer, which places no constraint on the intermediate thought itself. The paper identifies four failures that follow from this, and introduces REST, an objective that translates causality, minimality, separability, and stability into differentiable losses added to CE without architectural changes or added parameters at inference.

Key Contributions

  1. Identification of four failures of CE-only thoughts that lower the probability of the correct answer: substitution divergence, information loss, thought collision, and producer uncertainty. In trained systems, the paper reports that CE-only thoughts collapse across distinct questions and retain irrelevant information.
  2. The REST objective, which translates four theoretically motivated properties of thought representations into differentiable loss terms added to the CE objective, each weighted by a hyperparameter beta.
  3. An extension of latent recursive systems from the multi-agent setting to a single agent that recurs on its own hidden states, with the base LLMs and the inner link frozen and only the outer link (and the pooling probe for separability and stability) trained.
  4. An evaluation across 7 benchmarks spanning mathematics, science, medicine, and code generation, in both single-agent and multi-agent settings, under matched training data, compute, and latent budget.

Main Findings

  • Accuracy gains over CE-only: Averaged over all REST configurations, accuracy improves by +3.5 points in the multi-agent setting and +3.3 in the single-agent setting. The best configurations reach +7.5 and +6.5 points respectively, and the paper states gains of up to 7.5 percentage points across agent settings and model sizes.
  • Multi-agent results (Table 3): In the Light system every setting improves accuracy over CE-only. In the Scaled system this holds except for separability, which lowers accuracy by 2.5 points and reduces tokens by 9.3%. The largest average change is Best pair in the Scaled system at +7.5, followed by Minimality in the Scaled system at +6.9 and All properties in the Scaled system at +6.8.
  • Single-agent results (Table 2): Every property term improves accuracy over CE-only. Minimality gives the largest single-property gain: +3.0 points in Light and +6.5 in Scaled. Stability in Scaled reaches +6.4, and Causality in Scaled +4.8. "All properties" gives +2.9 in Light and +1.4 in Scaled; "Best pair" gives +1.0 in Light and +4.1 in Scaled.
  • Which property helps depends on the consumer: Causality leads when the thought passes to a different agent, while minimality leads when a model recurs on its own thought.
  • Convergence on a final answer: REST reaches a 95% boxed-answer rate compared with 73% for CE, which the paper describes as a 30% improvement in convergence on a final answer.
  • Token cost: REST decodes 15.4% more tokens on average at inference. In the multi-agent Light system, REST increases tokens (+11.5% to +29.9%); in the single-agent Light system, token counts are lower than CE for several configurations, from -2.9% to +1.6%. In the Scaled configurations, changes range up to +47.5% in the single-agent setting and +37.9% in the multi-agent setting.
  • Minimality in the single-agent Light system exceeds the baseline on every task while decoding fewer tokens overall, indicating its accuracy gain does not come from longer generation.
  • Against auxiliary-loss baselines (Table 4): REST averages +3.8 (Light) and +3.1 (Scaled). CODI (beta = 20) reaches +3.3 (Light) and +0.1 (Scaled); SIM-CoT (lambda_step = 0.3) reaches +3.2 (Light) and -4.4 (Scaled). REST also retains an advantage in token overhead on the Light system.
  • Thought content: Under REST, a thought encodes more of the producer's own output (its plan or refined plan) than of the prompt it received; CE encodes relatively more content from the input. The effect grows further along the chain of latent thoughts.
  • Preserved information: REST recovers more of the accuracy obtained when the solver is given the oracle refined plan instead of the thought.
  • Collapse: A PCA projection of the thought for 200 random samples shows CE thoughts gathered into dense clusters, which the paper reports as collapse with no shared content; REST spreads thoughts apart.
  • Superposition: Measured with a metric adapted from Deng et al. (2026), REST maintains effective superposition and slightly increases it as the weight beta increases.
  • Recursion depth: Averaged over CE-only, additional rounds increase REST's gain for the Scaled agents and reduce it for the Light agents, in both the single-agent and the multi-agent systems. Larger agents therefore use additional rounds of latent exchange; smaller agents obtain their gain from a single round.
  • Weight sensitivity: The r = 1 sweep in Table 11 shows both causality and minimality improve average accuracy at every beta evaluated, so the result does not rely on careful weight tuning.
  • Training CE under causality (Figure 7(b)): CE-only plateaus above causality for Light, where it does not memorize, while for Scaled it memorizes the training set and still underperforms REST in accuracy.

Methodology in Plain English

The authors study the recursive latent multi-agent system of Zou et al. (2026a) and extend Definition 2.1's transfer operation to a single agent. In the multi-agent setup, three frozen agents (a planner, a refiner, and a solver) each take a latent step budget of m' positions; the planner proposes an approach, the refiner revises it, and the solver produces the answer. Hidden states pass through an inner link (frozen, mapping a hidden state back into the agent's own input space) and a trained outer link (mapping a producer's output into a consumer's input space), and the whole chain loops for several rounds.

REST keeps the agents and inner link frozen and trains only the outer link, adding one beta-weighted term per property at every transfer on top of the CE of the final decoded answer. Causality penalizes the KL divergence between the consumer's next-token distribution under the producer's text embeddings and under the thought. Minimality is a difference between two cross-entropy terms: the consumer's CE on the producer's output given the thought, minus its CE on the producer's input given the output and the thought, so that input content irrelevant to the output is penalized. Separability pools each thought into a single vector with an attention pool and penalizes cosine similarity to the K preceding thoughts in the batch with a temperature-scaled log-sum-exp. Stability trains a probe over the thought to predict the producer's mean predictive entropy along its output.

Training uses the Sequential-Math dataset built from s1K and m1K question-answer pairs rewritten into role-specific texts, with AdamW under a cosine learning rate schedule, three training seeds, and only the outer link and pooling parameters updated. Evaluation runs on the Light system (Planner Qwen3-1.7B, Refiner Llama-3.2-1B-Instruct, Solver Qwen2.5-Math-1.5B-Instruct) and the Scaled system (Planner Gemma-3-4B-it, Refiner Llama-3.2-3B-Instruct, Solver Qwen3.5-4B), against CE-only, CODI, SIM-CoT, a frozen-LLM text baseline, Best pair, and All properties.

Why This Matters

Impact on research. The paper argues that final-answer CE gives an incomplete view of latent reasoning systems: the collapse and input retention it documents are invisible in the CE objective, so inspecting the thought representation should accompany accuracy reporting. It also makes persistent or latent communication easier to decode and audit, since REST thoughts contain relatively more of the producer's output than of its prompt.

Real-world applications (based on the benchmarks the paper evaluates):

  • Competition and exam-style mathematics, evaluated on MATH500, AIME2025, and AIME2026.
  • Scientific question answering, evaluated on GPQA-Diamond.
  • Medical question answering, evaluated on MedQA.
  • Code generation, evaluated on MBPP+ and LiveCodeBench-v6, with a lower inference temperature used for code than for other reasoning tasks.

Industry relevance. The method requires no architectural change and adds no parameters at inference, so it can be added to existing latent recursive systems without modification; the paper reports that it trains only the outer link, which is a cheap component relative to the base models. The tradeoff is token cost: 15.4% more tokens on average at inference, attributed to a higher rate of converging on a final answer.

Future Directions

  • Extending REST beyond the outer link, since the paper trains only that link while the base LLMs and the inner link remain frozen, a choice made to isolate the contribution of each property term. Extending it to the inner link or to the agent itself is left to future work.
  • Sweeping all combinations of properties and hyperparameters, which the authors do not do because the number of required runs grows multiplicatively with each addition.
  • Replacing the cheap entropy-based approximation of Stability, opted for under compute limitations, with a more faithful estimate.
  • Determining how the best-performing property depends on the consumer, given that causality leads when a thought passes to another agent and minimality leads when a model recurs on its own thought.

Target Audience

Researchers and engineers working on latent reasoning, chain-of-thought alternatives, and multi-agent LLM systems, particularly those who already run or plan to run latent recursive architectures and need a training signal for the intermediate representation. It is also relevant to readers interested in evaluation methodology for latent systems, since it argues for inspecting thought representations alongside answer accuracy. The formal derivations in the appendix make it more suitable for readers comfortable with probabilistic notation than for newcomers.

Authors’ abstract

Large language models can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30\%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: https://fard-lab.github.io/REST

Read the original paper