Research
Dual-End Consistency Model
Overview Research area: Computer vision and generative modeling, specifically efficient sampling for diffusion and flow-based models via consistency model distillation. Technical level: Advanced. The

- arXiv
- 2602.10764
- Published
- 2026-02-11
- Authors
- Linwei Dong, Ruoyu Guo, Ge Bai, Zehuan Yuan, Yawei Luo, Changqing Zou
AI summary
Overview
Research area: Computer vision and generative modeling, specifically efficient sampling for diffusion and flow-based models via consistency model distillation.
Technical level: Advanced. The paper assumes familiarity with probability flow ODEs, flow matching, consistency models, Jacobian-vector products (JVPs), and total-variation error bounds.
Scope in one sentence: The paper diagnoses why consistency models train unstably and sample inflexibly, then proposes a distillation framework (DE-CM) that optimizes three selected sub-trajectory clusters to reach a reported 1.70 FID at one-step generation on ImageNet 256×256.
Author affiliations listed: Zhejiang University, Bytedance Inc., and Zhejiang Lab. The paper is listed as arXiv:2602.10764v3 [cs.CV], dated 28 Jun 2026.
What This Paper Is About
Diffusion and flow-based generative models produce high-quality images but require many slow iterative sampling steps, which limits practical deployment. Consistency models (CMs) are a leading distillation approach for few-step generation, yet they suffer from training instability and inflexible sampling — the model is locked to a fixed step count, accumulates artifacts in multi-step sampling, and is sensitive to the initial noise.
This paper argues that the root cause both problems share is a poor choice of optimization trajectory, and proposes the Dual-End Consistency Model (DE-CM), which optimizes only three carefully selected trajectory clusters instead of the whole trajectory space.
Key Contributions
-
An analysis of CM failure modes. The authors decompose the continuous-time CM training objective and show that training instability comes from loss divergence between a supervised term and an unstable self-supervised term, while sampling inflexibility comes from error accumulation in the CM sampler (which operates in the γ=1 regime).
-
DE-CM, a sub-trajectory selection framework. Rather than optimizing all sub-trajectories (as MeanFlow and AYF do, which the authors characterize as O(n²) complexity and prone to entangled objectives), DE-CM selects three trajectory clusters: the consistency trajectory ([t,1]), the instantaneous trajectory (s=t), and the noise-to-noisy (N2N) trajectory.
-
A new noise-to-noisy (N2N) mapping objective. This objective maps pure noise to any intermediate noisy state rather than only to data x₁, which the authors state alleviates error accumulation from long-jump CM sampling. They derive the gradient convergence of this objective and note that N2N differs from BOOT, which is a data-free bootstrapping method based on a discrete consistency property.
-
Reported state-of-the-art one-step result. DE-CM achieves a reported FID of 1.70 with 1 NFE on ImageNet 256×256 and 1.33 with 2 NFE, using 250 training epochs and a 675M-parameter model.
Main Findings
-
One-step ImageNet result: DE-CM reports FID 1.70 at 1 NFE with 675M parameters and 0.011 time cost, and FID 1.33 at 2 NFE (0.024 time). At 50 NFE it reports FID 1.26 with 0.58 time. The paper states this is a new state of the art at 1 NFE among comparable-scale models, and that 2 NFE surpasses the vast majority of multi-step models.
-
Comparison against CM baselines (1 NFE / 2 NFE FID): MeanFlow 3.43 / 2.20; MeanFlow-† (distillation reproduction under the authors' 250-epoch configuration) 2.79 / 1.74; IMM 7.77 / 3.99; FACM-† 1.76 at 1 NFE; iCT 34.24 / 20.30; sCMs-† 3.34 / 1.94; Shortcut 10.60 at 1 NFE.
-
Comparison against GANs, masked/AR, and diffusion baselines: GigaGAN 1 NFE, 569M, FID 3.45; StyleGAN-XL 1 NFE, 166M, FID 2.30, 0.3 time; VQGAN 256 NFE, 227M, FID 3.04; MaskGIT 8 NFE, 227M, FID 6.18; VAR-d30 10 NFE, 2B, FID 1.92; ADM 250 NFE, 280M, FID 10.94; DiT-XL/2 (w=1.5) 250×2 NFE, 675M, FID 2.27; SiT-XL/2 (w=1.5) 250×2, 675M, FID 2.15; LightningDiT-XL/1 (w=6.7) 250×2, 675M, FID 1.35. The paper notes DE-CM's 1.70 does not beat the LightningDiT teacher's 1.35 multi-step score.
-
Text-to-image feedback scores (CLIP / BLIP / ImageReward): DE-CM at 1 NFE 0.2996 / 0.5398 / 0.5758; at 4 NFE 0.2999 / 0.5474 / 0.8117; at 8 NFE 0.3000 / 0.5479 / 0.8671; at 50 NFE 0.2993 / 0.5490 / 0.9712. For context, Flux.1-Dev at 50×2 NFE reports 0.2941 / 0.5384 / 0.9634, and SD3.5-Medium at 50×2 reports 0.2998 / 0.5451 / 0.8750. The paper highlights BLIP scores as notably strong and states the model performs well across all NFE settings.
-
Ablation across inference steps (FID at STEP=1 / 2 / 4 / 50): Baseline Teacher 309.51 / 230.79 / 111.59 / 2.77; Baseline Teacher (w=1.75) 262.60 / 221.03 / 87.56 / 1.59; Flow Matching only 263.86 / 220.55 / 87.93 / 1.45; CM distillation only 3.08 / 2.50 / 2.54 / 6.52; FM + CD 1.78 / 1.41 / 1.62 / 1.78; FM + N2N 441.66 / 297.28 / 132.60 / 1.41; FM + CD + N2N (full DE-CM) 1.70 / 1.33 / 1.41 / 1.26. Removing CD prevents few-step inference; removing FM hurts quality through unstable training; removing N2N hurts multi-step sampling through loss of "initial noise perception."
-
Gradient stability finding: The authors train the supervised and self-supervised components of the decomposed CM objective independently and report that the self-supervised loss is more prone to instability and volatility than the fully supervised term. They attribute this to divergence when F_θ has not converged and to the consistency term overshadowing the pretrained instantaneous velocity concept.
-
Sampling error bound: Citing the consistency trajectory model analysis, the authors state that with the interpolation parameter γ=1, approximation error aggregates to O(√(T + t₁ + ... + t_N)), while γ=0 gives a stable bound of O(√T). CM sampling is constrained to γ=1.
-
Efficiency claims: DE-CM is reported to converge efficiently at 16 GPU hours under 8 GPU resources, with clearer images than MeanFlow and fewer artifacts than sCMs in 1 NFE visual comparisons.
-
Related-method comparison (Table 3): AlphaFlow 1NFE-FID 2.58, 2NFE-FID 2.15, and the authors' reproduction 2.24 / 1.71; CMT 3.34 / not reported, reproduction 2.71 / 1.68; DE-CM only reproduction numbers 1.70 / 1.33.
Methodology in Plain English
The authors start by asking why consistency models are hard to train and why they sample poorly. They rewrite the CM training loss as two pieces: one that is supervised by the true velocity from the pretrained model, and one that is self-supervised (the model compares itself against a frozen copy of itself). They show that the self-supervised piece produces much less stable gradients and that it drifts away from the pretrained model's velocity knowledge.
On the sampling side, they explain the CM sampler effectively takes long jumps straight to the data at every step, so errors overlap and compound as step count grows. The model also has no experience of intermediate noise levels during training, since it is always asked to jump all the way to x₁.
Their fix is to stop optimizing every possible sub-trajectory and instead pick three:
-
Consistency trajectory (s = 1). Train the model to map any noisy point to the clean data, which is what gives few-step distillation. This reuses continuous-time consistency distillation and computes the needed derivative with a Jacobian-vector product (JVP).
-
Instantaneous trajectory (s = t). When the start and end time are the same, the self-supervised term vanishes and the loss is driven solely by the pretrained teacher's velocity — a deterministic, supervised objective. The authors use this as a boundary regularizer (combined with an additional cosine similarity term) to stabilize training and preserve the pretrained model's ability to do multi-step ODE sampling.
-
Noise-to-noisy trajectory (r → 0, target t). Train the model to map pure noise to an arbitrary intermediate noisy state, not just to data. This exposes the model to inputs resembling what it sees at inference and reduces error accumulation.
Training combines all three losses with a weighting function w(r,t) and hyperparameters λ and γ. The pipeline uses a pretrained teacher (LightningDiT for class-to-image, with classifier-free guidance applied through a w_cfg term), an online model initialized from the teacher, and an EMA copy. For class-to-image they train in a pretrained VAE latent space of 32×16×16 with learning rate 1e-4 and the AdamW optimizer, evaluating FID on 50K generated images. For text-to-image they use 100K samples from the text-to-image-2M dataset, the default SD3-VAE tokenizer, LoRA with rank 64, learning rate 5e-4, and AdamW.
Why This Matters
Impact on research. The paper reframes CM instability as a trajectory-selection problem rather than an architecture or regularization problem, and provides a decomposition of the continuous-time objective to support that framing. If the analysis holds, it gives the distillation community a more principled criterion for choosing training targets — and it distinguishes DE-CM from flow-map methods (MeanFlow, AYF), Shortcut-style methods, and BOOT in both motivation and formulation. The claimed benefit of stable performance from a single set of parameters across NFE budgets (1 through 50) also bears on the common complaint that distillation methods trade multi-step quality for few-step speed.
Potential real-world applications (grounded in domains the paper mentions as benefiting from diffusion and flow models — image, 3D, audio, video, and text-to-image generation):
- Interactive image and text-to-image generation where sub-second response matters; the paper reports a 0.011 time cost at 1 NFE.
- On-device or edge generation, where a fixed budget of 1-2 network evaluations reduces compute and energy needs.
- Video and audio generation pipelines, where per-frame or per-segment sampling cost multiplies quickly.
- 3D asset generation, where iterative sampling must fit into interactive design workflows.
Industry relevance. The reported 675M parameter model and 250-epoch training schedule, plus the LoRA-based text-to-image setup with the SD3.5 and Flux model families as baselines, positions this work in the territory of deployable production image generators. Faster sampling directly reduces inference cost per image, which is a primary economic driver for generative image services.
Future Directions
-
Resolving the stated limitation. The paper's conclusion begins to note a limitation related to "the incompatibility of the JVP opera[tion]" but the provided text is truncated, so the full limitation and its proposed resolution are not reported here. Completing or working around that incompatibility is an obvious next step.
-
Extending DE-CM beyond image generation. The paper demonstrates class-to-image and text-to-image only; whether the three-trajectory recipe transfers to 3D, audio, or video generation is not tested.
-
Scaling the training budget. Results are reported with 250 training epochs; behavior at longer schedules or larger model sizes is not reported.
-
Broadening the comparison set. The text-to-image comparisons use 100K training samples from text-to-image-2M, and the class-to-image comparisons use 675M-parameter backbones; whether DE-CM's advantage persists with other teachers, larger data, or non-LoRA fine-tuning is an open question.
Target Audience
This paper is most useful to generative-model researchers and engineers working on distillation and few-step sampling — particularly those already familiar with consistency models, flow matching, and PF-ODE solvers. It will also interest practitioners who need to deploy diffusion or flow models under tight inference budgets, and readers tracking the ongoing competition among flow-map, shortcut, and consistency-based distillation families on the ImageNet 256×256 benchmark. Beginners without a background in stochastic differential equations and ODE-based generative modeling will find the analysis sections difficult, since the derivations rely on gradient convergence arguments and Jacobian-vector products.
Authors’ abstract
The slow iterative sampling nature remains a major bottleneck for the practical deployment of diffusion and flow-based generative models. While consistency models (CMs) represent a state-of-the-art distillation-based approach for efficient generation, their large-scale application is still limited by two key issues: training instability and inflexible sampling. Existing methods seek to mitigate these problems through architectural adjustments or regularized objectives, yet overlook the critical reliance on trajectory selection. In this work, we first conduct an analysis on these two limitations: training instability originates from loss divergence induced by unstable self-supervised term, whereas sampling inflexibility arises from error accumulation. Based on these insights and analysis, we propose the Dual-End Consistency Model (DE-CM) that selects vital sub-trajectory clusters to achieve stable and effective training. DE-CM decomposes the PF-ODE trajectory and selects three critical sub-trajectories as optimization targets. Specifically, our approach leverages continuous-time CMs objectives to achieve few-step distillation and utilizes flow matching as a boundary regularizer to stabilize the training process. Furthermore, we propose a novel noise-to-noisy (N2N) mapping that can map noise to any point, thereby alleviating the error accumulation in the first step. Extensive experimental results show the effectiveness of our method: it achieves a state-of-the-art FID score of 1.70 in one-step generation on the ImageNet 256x256 dataset, outperforming existing CM-based one-step approaches.