Research
Scaling and Distilling Text Embeddings for Better Diffusibility
Overview Research area: Natural Language Processing — specifically continuous diffusion language models (DLMs) and the text embedding spaces they operate on. Technical level: Advanced. The paper assum

- arXiv
- 2610.01016
- Published
- 2026-10-01
- Authors
- Zekai Zhang, Yunjie Tian, Yanjin He, Xiaoyan Zhang, Dongdi Zhao, Qing Qu, Di Fu
AI summary
Overview
- Research area: Natural Language Processing — specifically continuous diffusion language models (DLMs) and the text embedding spaces they operate on.
- Technical level: Advanced. The paper assumes familiarity with diffusion models, latent/continuous diffusion, encoder-decoder architectures, and knowledge distillation, though the core ideas are stated in accessible terms.
- Scope in one sentence: The paper isolates the choice of text embedding as the single variable in a continuous DLM and tests whether scaling to stronger embedding models, and then distilling the strongest one into a smaller student encoder, produces a latent space that is easier for diffusion to generate.
What This Paper Is About
Continuous diffusion language models generate text by denoising continuous text embeddings and then decoding them back into tokens, so the quality of the embedding space is central to whether they work at all. The paper asks a practical question the authors phrase as: which embedding makes the best latent space, i.e., the most diffusible? It holds the diffusion machinery fixed (the ELF framework) and changes only the embedding model, first by scaling up to stronger members of the T5/T5Gemma family, then by distilling the strongest one into a smaller student encoder that learns to place plausible alternative words closer together.
Key Contributions
- Scaling embeddings as a route to diffusibility. The authors show that moving to stronger embedding models within the same family (T5-small/base to T5Gemma-1 to T5Gemma-2) makes a better latent space for continuous DLMs, establishing a strong baseline with T5Gemma-2 embeddings.
- A definition and measurement of diffusibility. They study diffusibility from the perspective of generating valid, decodable embeddings, defining an embedding as invalid when no candidate word lies within reach, and quantify how often standard embeddings fail this test.
- Distilling decoder probabilities for better geometry. They enhance T5Gemma-2's diffusibility by distilling its decoder's soft-label probabilities into a student encoder, showing that this trades some discriminative power for generation quality.
- A controlled evaluation across scales and corpora. The distilled embeddings are evaluated with the same ELF-B and ELF-M diffusion models on OWT-1024 and LM1B, together with a cost breakdown of scaling the embedding side.
Main Findings
- Scaling improves diffusion quality. Along T5-small → T5Gemma-1 → T5Gemma-2, diffusion performance improves, and T5Gemma-2 cuts Gen. PPL by about 40% at the same entropy compared with the T5-small embedding used in ELF. T5Gemma-2 was the only embedding space tested that reached real-text entropy.
- Being strong on discriminative tasks is not the same as being diffusible. ModernBERT, a strong modern encoder-only model, performed well on downstream tasks but was less suitable for diffusion and generated repetitive sentences.
- Scaled embeddings have a failure mode. T5Gemma-2 embeddings keep plausible alternative words for a position apart, so an imperfect sampling trajectory can end at an embedding that is near none of them—described by the authors as an invalid embedding that decodes with low confidence or to a wrong word.
- Quantified failure rates. Using a (0.91) nearest-neighbor threshold for invalidity and a top-1 probability below (0.6) for uncertainty, over 64 generated sequences of 1024 tokens at NFE (32), the teacher has (142) uncertain positions per sequence, (109) of which are also invalid; the student has (42) and (12).
- Distillation fixes the geometry. Distilling the teacher decoder's probabilities makes the student pull alternative embeddings closer, forming a more connected and diffusible latent space; the teacher's uncertain embeddings land far (mean (1.39) nn, median (1.38)), the student's land close (mean (0.82), median (0.81)).
- Soft labels are the active ingredient. Replacing KL on soft labels with cross-entropy on one-hot input tokens produced students that never reached real-text entropy (entropy (\leq 4.18)), and MSE on the teacher's embeddings also failed to reach it ((\leq 5.26)); a randomly initialized student cannot converge.
- Student depth is not the source of the gain. An 18-layer student, as deep as the teacher, still beats the teacher, at entropy (5.05) reading (31)–(34) Gen. PPL for 12–18 layers versus (49) for the 9-layer student and (46) for the teacher; the authors use 9 layers to halve the encoding cost.
- Best headline result. On OWT-1024, the medium-sized DLM (Student + ELF-M) reaches Gen. PPL (17.8 \pm 0.1) at entropy (5.45 \pm 0.01) and MAUVE (0.89 \pm 0.01), against real-text PPL (15.4 \pm 0.3), outperforming GPT-2-M on Gen. PPL (GPT-2-M: (20.8 \pm 0.3)). The same Student + ELF-B reaches Gen. PPL (31.2 \pm 0.2) at entropy (5.44 \pm 0.00), versus (38.5 \pm 0.5) for the T5Gemma-2 teacher.
- Improvement generalizes to another corpus. On LM1B, Student + ELF-B reaches Gen. PPL (50.5 \pm 0.5) at entropy (4.31 \pm 0.00), versus (60.6 \pm 0.3) for T5Gemma-2 + ELF-B.
- More stable sampling. T5Gemma-2 collapses at extremely low NFEs such as 16–32 due to under-integration, and at NFE 64 can still collapse to a language mix, while the distilled embeddings are more stable; improvement holds at every NFE tested from 64 to 512.
- A discriminative cost. The student degrades on downstream classification ((89.4) to (78.1) on SST-2, and lower on STS-B and MRPC), but keeps round-trip reconstruction strong (top-1 token accuracy (99.3) for both teacher and student on WikiText-103, and (99.3) vs. (99.4) on held-out arXiv papers).
- Scaling embeddings is cheap on the diffusion side. Encoding is frozen and can be cached; the distilled student halves the encoding cost of T5Gemma-2 (Student: (50_{+168}) M params, (0.10) T FLOPs, (2.1) ms per 1024-token sequence, versus (100_{+168}) M, (0.21) T, (3.4) ms). End-to-end, ELF-M produces a 1024-token sequence in (1.5) s at NFE 64 and (11.8) s at NFE 512, while GPT-2-M takes (10.7) s with 1024 serial steps.
Methodology in Plain English
The researchers took an existing continuous diffusion framework, ELF, and deliberately left nearly all of it alone—training recipe, sampling, noise scheduling, self-conditioning. The only thing they changed was the frozen embedding model that maps tokens into the latent space the diffusion model denoises (plus the final LM head, since vocabularies differ).
First, they ran the same small diffusion model on a 25% subset of OpenWebText using 512-token sequences, comparing embeddings from T5-small, T5-base, T5Gemma-1 variants, T5Gemma-2-270M, and ModernBERT, plotting generative perplexity against entropy to see which latent space reached the entropy of real text at the best quality.
Second, having identified T5Gemma-2 as the best of these, they inspected generated embeddings and found a structural problem: because the model was trained to reconstruct its input against one-hot targets, plausible alternatives for a token position sit far apart in embedding space, so a noisy diffusion trajectory can end up nowhere near any of them. To diagnose this they took, for each position, the words the deployed decoding head would produce as candidates, re-encoded the sequence with each candidate in place, and measured distances in units of the median nearest-neighbor distance between real embeddings.
Third, they trained a student encoder to imitate the teacher decoder's probability distribution rather than its hard target—a KL divergence over the teacher's top-256 logits (renormalized) from a 262k vocabulary, with the teacher encoder and decoder frozen. The student was initialized from 9 evenly spaced layers of the teacher's 18 layers and trained on the same packed 1024-token OpenWebText sequences. Finally, they retrained the same ELF-B and ELF-M diffusion models on these student embeddings and compared generation quality on OWT-1024 and LM1B, re-tokenizing samples with the GPT-2 tokenizer for comparability and reporting the best Gen. PPL under GPT-2-Large at or above real-text entropy, averaged over 3×1024 samples from 3 seeds.
Why This Matters
Impact on research. The paper argues that designing the embedding space of a continuous DLM is as important as designing the diffusion process itself. It gives the text counterpart of a finding already established in image diffusion, where replacing VAE latents with pretrained representations improves diffusibility, and it identifies a tension between discrimination and generation: embeddings that are excellent for classification can be poor latent spaces for diffusion. It also offers a concrete measurement vocabulary—valid, uncertain, invalid embeddings—for a failure mode that prior work mostly described qualitatively.
- Faster parallel text generation. Continuous DLMs denoise rather than emit tokens serially; the paper reports ELF-M producing a 1024-token sequence in 1.5 s at NFE 64 versus 10.7 s for GPT-2-M, making this route relevant for latency-sensitive generation.
- Compact deployment of strong representations. The distillation halves the encoding cost of T5Gemma-2, which matters where a large encoder is the bottleneck at inference time.
- Question answering and other practical NLP tasks. The authors explicitly name practical-scale models for real-world tasks such as QA as a next step, framing this work as a basis for that.
- Reuse of existing pretrained encoders. Because T5Gemma-2 is itself adapted from Gemma-3, the approach suggests a path for turning already-trained language models into latent spaces for diffusion rather than training embedding models from scratch.
Industry relevance. The paper reports an explicit cost decomposition (parameters, FLOPs, latency per 1024-token sequence on one H100 in bf16) for scaling the embedding side, showing that encoding is a small fraction of training and sampling cost and can be cached. That, plus a distillation recipe that cuts encoder size roughly in half, is the kind of accounting an engineering team needs when deciding whether the embedding or the diffusion trunk is the right place to spend compute.
Future Directions
- Few-step or one-step generation. The authors note that ELF trained on raw T5Gemma-2 under-integrates at low NFE, and name few/one-step generation as an open problem.
- Scaling to larger embedding models and practical tasks. They deliberately kept to smaller embedding models for controllable experiment size, leaving scaling to bulkier encoders (for example T5Gemma-2-1B/4B) and practical-scale models for real-world tasks such as QA to future work.
- Adapting autoregressive foundation models directly into embeddings. Since T5Gemma-2 is adapted from Gemma-3, they suggest skipping that
Authors’ abstract
Diffusion language models (DLMs) offer a promising alternative to autoregressive (AR) language generation. Recent advances in continuous DLMs, which apply latent diffusion to continuous text embeddings, raise a practical question: which embedding makes the best latent space, i.e., the most diffusible? To answer this, we search through different embeddings and find that scaling the embedding model to stronger ones within the same family (T5 to T5Gemma-1 to T5Gemma-2) greatly improves generative performance. But the raw T5Gemma-2 embeddings are still not optimal. They are so discriminative that even the embeddings of plausible alternative words are separated, which makes the generation vulnerable to imperfect sampling. Consequently, continuous diffusion often fails to reach any of them and ends up at an invalid embedding instead. To address this, we distill T5Gemma-2 into a student encoder that learns the teacher's decoded probabilities as soft labels. Learning from such soft labels makes the student pull the alternative embeddings closer while maintaining the encoding-decoding mechanism. The distilled embeddings form a more connected and diffusible latent space, improving over the vanilla T5Gemma-2 embeddings. As a result, our medium-sized DLM achieves Gen. PPL 17.8 (against real-text PPL 15.4) at real-text entropy on OpenWebText, outperforming GPT-2-M on Gen. PPL.