Research
What Matters for Latent Reasoning with Flow Matching
Overview Research area: Latent reasoning in large language models, specifically generating intermediate reasoning in a continuous latent space using flow matching (the same family of techniques behind

- arXiv
- 2610.06666
- Published
- 2026-10-05
- Authors
- Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos
AI summary
Overview
Research area: Latent reasoning in large language models, specifically generating intermediate reasoning in a continuous latent space using flow matching (the same family of techniques behind latent diffusion models), rather than writing out an explicit chain of thought (CoT) token by token.
Technical level: Advanced. The paper assumes familiarity with variational autoencoders, flow matching / diffusion, classifier-free guidance, and LLM fine-tuning.
One-sentence scope: The paper defines five requirements for an effective latent thought, identifies through controlled ablations the training choices that make latent flow matching work for reasoning, and packages them into a recipe called Flow-based Latent Reasoning (FLaRe).
What This Paper Is About
Explicit chain-of-thought reasoning is effective but costly: its cost grows with chain length, much of that cost is spent on prose that adds little to the reasoning, and a single rollout commits to its early tokens, so exploring another path requires more generation. Latent reasoning is meant to fix this by carrying intermediate computation in continuous states and verbalizing only the answer.
The problem the authors address is that existing latent methods largely fail to deliver on that promise — they learn shortcuts from the question, distill the explicit CoT into their weights, or imitate it one token at a time. The paper's goal is to pin down, through ablation, exactly which training decisions make latent flow matching genuinely useful for reasoning, and to combine them into one recipe.
Key Contributions
-
A simple training recipe. Through controlled ablations spanning the latent space, the flow training, the answer readout and a final self-training stage, the authors identify the choices that make latent flow matching work for reasoning and combine them into FLaRe.
-
Five requirements for latent reasoning, and a probe for each. The paper defines what an effective latent thought must be — useful, diverse, explainable, refinable and efficient — designs a targeted probe for each requirement, and reports that FLaRe improves on prior latent methods (Coconut, CODI, PCCoT) on all five.
-
Improved accuracy and efficiency. FLaRe compares favorably with prior latent methods on arithmetic benchmarks and reaches 97% of explicit CoT accuracy at a quarter of its latency.
-
An end-to-end two-stage pipeline. Stage 1 trains a shared flow model to denoise latent codes and answer from its own thoughts; stage 2 self-trains the model on its own verified rollouts for new questions using only questions and reference answers, with the answer loss backpropagated through the full denoising path.
Main Findings
-
Reconstruction quality is not the goal. Across 66 VAEs from nine sweeps, reconstruction and decoded accuracy correlate at −0.12. At dimension d = 512, growing the code from 8 to 32 slots raises reconstruction from 96.8% to 99.4% but drops direct accuracy from 47.1% to 25.4% and decoded accuracy from 53.2% to 32.3%. Every detail the code preserves is one more detail the flow must generate.
-
Smoothness must be engineered. With β fixed at 10⁻⁵, the authors corrupt the encoder input and latent code during VAE training: token substitution with probability p_sub = 0.3, variance-preserving noise with δ = 0.7 on half the sampled codes, and latent dropout with p_drop = 0.4. All three together give 47.1 direct and 53.2 decoded, better than removing any one of them (e.g. removing token substitution gives 45.6 direct and 47.6 decoded).
-
Symbolic CoTs are the better target than natural language CoTs. Encoding the natural language CoT cost about 10 direct and 17 decoded points versus the default. Symbolic input with dual decoder routes gives 50.7 direct / 61.3 decoded, versus 40.9 / 44.7 for natural language input on a language-only route.
-
A stronger decoder helps downstream accuracy. With the 1B encoder fixed, a fine-tuned 3B decoder gave 47.1 direct / 53.2 decoded, the best among the decoder sizes and training choices tested (1B frozen: 37.2 / 43.4; 1B fine-tuned: 40.8 / 46.7; 3B frozen: 41.9 / 47.7).
-
The flow should be trained mostly near noise. Uniform time performed worst; performance improved only when the time distribution was strongly shifted toward noise. The final setting is a logit-normal distribution, t = sigmoid(s) with s ~ N(−2, 0.8²).
-
The answer reader must be trained on imperfect thoughts. Training only on exact codes gives 45.7 direct; a 50/50 mixture of noised codes and the model's own detached one-step endpoints gives 53.3 direct / 57.3 decoded. Letting gradients flow through those endpoints instead performs worse (50.2 direct / 56.6 decoded).
-
Several CoTs per question give the largest single data gain. Adding diverse CoTs to the flow training data raised the result from 51.9 / 56.7 to 54.0 / 59.7, while adding OOD data lowered it to 46.3 / 52.8.
-
Stage 2 helps, and its key ingredient is the full-rollout answer loss. Self-training on verified rollouts raised direct accuracy from 55.3 to 59.1 and decoded from 61.3 to 62.6, a gain of 3.8 and 1.3 points. Removing the answer loss (55.1 / 60.7), detaching the rollout endpoint (56.8 / 60.0) or reducing to a one-step rollout (56.3 / 60.7) loses most or all of the gain; verification is also needed (57.1 / 60.7).
-
Usefulness (twin test). Explicit CoT follows the twin's CoT on 97.2% of pairs, versus 39.7% for FLaRe stage 1, 37.3% for stage 2, 31.8% for CODI, 27.7% for PCCoT and 7.3% for Coconut. Restricting to pairs both models answered correctly before injection, stage 1 rises to 91% and stage 2 to 76%. The paper reports that Coconut's prior work found little change when thoughts are removed or swapped.
-
Diversity (sampling test). With 16 samples per question, baselines gain at most 5 points of coverage and majority voting never beats greedy. FLaRe stage 1 gains 16.6 points of coverage and 3.3 through voting; stage 2 gains 14.7 and 1.8.
-
Explainability (reading test). FLaRe's decoded thoughts contain all reference intermediate results in order on about half the questions, more often than any baseline. The baselines lack a dedicated thought decoder and must be read through the LM head, which the paper says recovers only fragments.
-
Refinability (budget sweep). Raising compute from 0.1 to 0.5 times explicit CoT adds at most 6.1 points to the baselines, and PCCoT loses 10.5 points. FLaRe's decoded reading gains 18.2 (stage 1) and 24.2 (stage 2) points, while its direct reading gains at most 1.8; the baselines peak at or below their training budgets.
-
Efficiency (latency test). CODI reaches 94% of explicit CoT accuracy at a 2.0× speedup and PCCoT reaches 91% at 3.9×. FLaRe stage 2, using the direct reading at two Euler steps, reaches 97% at 3.9×.
-
Benchmark comparison. On GSM8K, FLaRe stage 2 reaches 53.8 at Qwen2.5-0.5B-Instruct, 59.1 at Llama-3.2-1B-Instruct and 65.2 at Llama-3.2-3B-Instruct. It leads all latent methods on GSM8K at 0.5B and 1B, by 6.9 and 2.6 points over KaVa, performs comparably to KaVa at 3B, and leads or ties on GSM8K-Hard at every scale. SVAMP is the main exception, likely because nearly half of its questions add distractor numbers that are rare in GSM8K-Aug. Stage 2 improves every column at every scale, by 2.7 to 3.8 points on GSM8K, and stage 1 already gains 10.7 to 21.0 points on GSM8K over LaDiR, its most direct competitor.
Methodology in Plain English
The pipeline has three parts.
Compression. A variational autoencoder learns to squeeze an explicit chain of thought into a small latent code of 8 slots of dimension 512. The encoder reads a symbolic CoT — a version stripped of wording that records only the reasoning steps and intermediate results — followed by learned slot tokens. The decoder can read the code back two ways: as the symbolic CoT from the code alone, or, given the question, as the natural language version of the same reasoning. The authors deliberately make this latent space "smooth" by corrupting the input tokens and the sampled codes during training, so that a thought that lands near the right code still decodes to the same reasoning. They keep the KL weight small (β = 10⁻⁵) because increasing it degraded reconstruction before it improved generation.
Stage 1. A single shared flow model — Llama-3.2-1B-Instruct — learns two jobs at once. It denoises Gaussian noise into a latent thought conditioned on the question, and it generates the final answer from that thought. The flow loss is trained mostly at high noise levels, where inference starts, and the loss is averaged over four independent time/noise draws per code. The question is replaced by a null token with probability 0.1 so classifier-free guidance can be used at inference. A second forward pass teaches the model to read a 50/50 mixture of noised codes and its own detached one-step predictions, so the reader learns to cope with the imperfect thoughts it will actually see at test time. Inference runs 20 Euler steps at guidance scale 4.
Stage 2. For new questions that come with only answers, the model proposes 8 thoughts per question from pure noise over 20 guided steps. A frozen decoder turns each thought into a symbolic CoT, and only the distinct CoTs whose final answer matches the reference are kept and re-encoded as targets. Training then combines the flow loss on those targets, a replay of stage 1 data, and an answer loss computed on a fresh 10-step rollout whose gradients are backpropagated through every step and both guidance branches. This teaches the reader to use generated thoughts and the flow to produce thoughts that lead to the reference answer, without needing extra CoT annotations, a learned reward or a reinforcement learning objective.
Setup. The VAE encoder and decoder are initialized from Llama-3.2-1B and Llama-3.2-3B, pretrained on OpenMathInstruct-2 and then trained for 10 epochs on GSM8K-Aug (385K questions with symbolic CoTs).
Authors’ abstract
Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costing less than an explicit CoT at comparable accuracy. Current methods rarely meet these requirements: they learn shortcuts from the question, distill the explicit CoT into their weights, or imitate it one token at a time. We focus on flow matching in a learned latent space, the family we argue is best placed to meet them, and identify the training choices that make it work. The result is Flow-based Latent Reasoning (FLaRe), a simple recipe covering what the latent space encodes and how to shape it, where to train the flow, how to read out the answer, and a final stage of training on the model's own verified thoughts. A probe for each requirement shows that FLaRe improves on prior latent methods in all five. It also compares favorably with them on arithmetic benchmarks, while reaching 97% of the accuracy of explicit CoT at a quarter of its latency.