Skip to content
AI.info

Research

Unlocking the Potential of Diffusion Language Models through Template Infilling

Overview Research area: Natural Language Processing — inference-time conditioning strategies for Diffusion Language Models (DLMs). Technical level: Intermediate (requires familiarity with autoregressi

Unlocking the Potential of Diffusion Language Models through Template Infilling
arXiv
2510.13870
Published
2025-10-13
Authors
Junhoo Lee, Seungyeon Kim, Nojun Kwak

AI summary

Overview

Research area: Natural Language Processing — inference-time conditioning strategies for Diffusion Language Models (DLMs). Technical level: Intermediate (requires familiarity with autoregressive generation, masked diffusion, and standard NLP benchmarks). Scope: The paper proposes Template Infilling (TI), a training-free conditioning method that places structural anchors throughout the entire target response so a DLM fills masked segments between them, plus a Dynamic Segment Allocation (DSA) rule that expands segments when the model is uncertain.

What This Paper Is About

Diffusion Language Models can generate tokens at any position in any order, but existing inference methods still prompt them the way autoregressive models are prompted — by prepending context to the front of the sequence. The authors argue this throws away the model's bidirectional conditioning ability and causes instability, and they propose instead to embed a structural template (anchors plus masked spans) across the whole response so the model generates under global constraints.

Key Contributions

  1. Template Infilling (TI): A conditioning framework in which the prompt is deconstructed into a set of fixed anchors A = {A1, ..., An} distributed across the sequence, interleaved with masked spans M1 ... Mn that the model fills, so each span is conditioned on both preceding context and future anchors.
  2. Dynamic Segment Allocation (DSA): A protocol that monitors per-token confidence at each diffusion step and expands a mask segment by a fixed number of tokens δ when the least-confident token in that segment falls below a threshold τ, allowing the model to allocate more reasoning space where needed.
  3. Universality validation: The method is tested on two models covering different training paradigms — LLaDA-8B (trained from scratch with a diffusion objective) and Dream-7B (adapted from the autoregressive Qwen2.5-7B via fine-tuning), each in Base and Instruct variants.
  4. Evidence for structural (rather than prompt-engineering) gains: Ablations on template detail, anchor position, generation length, sampling steps, and a safety/jailbreak case study are used to argue that the improvement comes from global conditioning that induces a "System-2" reflective mode.

Main Findings

  • Average gain of 9.40%p: Across mathematical reasoning (GSM8K, MATH500), code generation (HumanEval), and trip planning (Constraint Satisfaction Rate), the paper reports a consistent average improvement of 9.40 percentage points over the baseline.
  • Prefix prompting does not help DLMs: In most rows of Table 1, prefix prompting gives negligible gains or degradation — for example, LLaDA-8B Base drops from 51.63 to 22.74 on GSM8K, and Dream-7B Base drops from 18.29 to 3.66 on HumanEval.
  • Largest single-model jumps: Dream-7B Base improves from 8.87 (Vanilla) to 44.58 with TI on GSM8K — described as nearly a five-fold improvement — and from 1.13 to 15.94 CSR on Trip Planning. LLaDA-8B Instruct improves from 49.58 to 71.49 on GSM8K.
  • One exception in the table: For LLaDA-8B Base, TI's average (26.26) is marginally below Vanilla (26.42), even though TI raises MATH500 (3.2 to 11.60) and HumanEval (35.4 to 28.05 is a decrease; MATH500 3.2 to 11.60 is an increase) — the paper's headline "consistent improvements" claim is stated over the baseline generally, not for every row.
  • DSA is the largest contributor: On GSM8K with Dream-7B, Vanilla scores 8.87 and Prefix Prompting 8.79 (−0.08). Static "Minimal" templates reach 24.94 (+16.07), detailed templates 36.00 (+27.13), and adding DSA reaches 44.58 (+35.71).
  • Robust to anchor position shifts: On GSM8K, TI Base scores 0.4458, with "Early" at 0.4033, "Late" at 0.4359, and "Compressed" at 0.4367 — the paper attributes the small variance to global conditioning rather than prompt tuning.
  • Stable under acceleration and long generation: With a fixed budget of 64 sampling steps, TI mitigates baseline degradation as generation length rises from 128 to 512 tokens, and holds accuracy better when sampling steps are reduced.
  • TI injects a sampling prior: In the Dream-Base model, whose generation order is otherwise chaotic, template anchors are generated first and regularize the sequence. The authors also report that Dream-Instruct reverts to a diagonal, autoregressive pattern, which they attribute to Dream's Context-Adaptive Noise Rescheduling and to instruction tuning leaving instruction tokens unmasked.
  • Safety behavior: In two jailbreaking scenarios (toxic substance generation, phishing script), the paper reports that autoregressive-style prompting and naive generation fail to refuse, while TI, using a "Draft-Critique-Refine" template embedded globally, produces a refusal.

Methodology in Plain English

Instead of pasting the prompt at the front and letting the model continue, the authors lay out the whole answer in advance as a skeleton: fixed anchor phrases (such as "Let me work through this problem" or "Step 1") separated by blank masked spans. Because a diffusion model can see the entire sequence at once, every blank is constrained by the anchors that come before it and the anchors that come after it, which the authors describe as boundary conditions that keep generation on a logical path.

Static blanks can be too short for hard problems, so DSA watches how confident the model is about each token in a blank. If the least confident token in a segment falls below a threshold, that segment grows by a fixed number of extra tokens, up to a maximum number of expansions. The position of subsequent anchors is preserved.

Evaluation is deliberately restricted to pure parallel generation, where the model must plan and produce all 128 tokens at once. A single static template per task is used (the same anchors for all math problems), with the DSA expansion rate capped at 8 tokens per step and at most 10 expansions. Both base and instruction-tuned variants of LLaDA-8B and Dream-7B are tested, using the authors' official codebases.

Why This Matters

The paper reframes the high degree of freedom in diffusion language models — usually treated as an instability to be suppressed with block-wise, semi-autoregressive schemes — as something to be exploited. It also shows gains on a model fine-tuned from an autoregressive backbone, suggesting the "diffusion property" can be acquired with relatively lightweight fine-tuning and that TI applies to that growing class of models.

Real-world applications implied by the benchmarks and demonstrations:

  • Step-by-step mathematical and competition-style problem solving (GSM8K, MATH500).
  • Code generation and fill-in-the-middle style program completion (HumanEval).
  • Multi-constraint trip planning, measured with Constraint Satisfaction Rate.
  • Safety-critical response workflows such as Draft-Critique-Refine refusal behavior.

Industry relevance: TI is training-free and compatible with existing DLM inference stacks, so it can be applied at inference time without retraining. Its reported stability under multi-token generation and reduced sampling steps is directly relevant to serving DLMs quickly and cheaply.

Future Directions

  • Train models with template objectives: The authors identify as a key limitation that current instruction-tuned models are still trained under the traditional prompt-inference paradigm, and propose incorporating TI into instruction fine-tuning itself.
  • Autonomous template generation: The paper envisions systems that synthesize a query-specific structural template, drawing inspiration from GEPA and evolutionary heuristics, turning structural guidance from a predefined constraint into an adaptive, self-generated blueprint.
  • Reconciling instruction tuning with global planning: Since Dream-Instruct reportedly reverts to a diagonal autoregressive generation pattern, an open question is how to preserve the global planning ability unlocked in the Base model while retaining instruction-following.
  • Beyond static templates: The current work uses static templates to establish feasibility; whether learned or dynamically adapted templates yield further gains is left open.

Target Audience

Researchers and engineers working on diffusion language models, non-autoregressive text generation, and inference-time conditioning; practitioners interested in constrained decoding, structured generation, or efficient parallel decoding under reduced sampling budgets; and safety researchers studying whether structural templates can enforce reflective workflows that soft prompts cannot guarantee.

Authors’ abstract

Diffusion Language Models (DLMs) have emerged as a promising alternative to Autoregressive Language Models, yet their inference strategies remain limited to prefix-based prompting inherited from the autoregressive paradigm. In this paper, we propose Template Infilling (TI), a tailored conditioning methodology for DLMs. Unlike conventional prefix prompting, TI flexibly aligns structural anchors across the entire target response space, establishing a global blueprint before filling in the masked segments. We demonstrate the effectiveness of our approach on diverse benchmarks, including mathematical reasoning, code generation, and trip planning, achieving consistent improvements of 9.40% over the baseline. Furthermore, we observe that TI provides additional advantages in multi-token generation settings, enabling effective speedup while maintaining generation quality and robustness. By enforcing these global constraints, TI ultimately facilitates System-2 reasoning, empowering the model to deliberate within a structurally defined solution space.

Read the original paper