Skip to content
AI.info

Research

Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation

Overview Research area: Efficient generative modeling, specifically compression and sparse restructuring of diffusion transformers (DiTs) for text-to-image synthesis. Technical level: Advanced. The pa

Dense2MoE: Restructuring Diffusion Transformer to MoE for Efficient Text-to-Image Generation
arXiv
2510.09094
Published
2025-10-10
Authors
Youwei Zheng, Yuxi Ren, Xin Xia, Xuefeng Xiao, Xiaohua Xie

AI summary

Overview

Research area: Efficient generative modeling, specifically compression and sparse restructuring of diffusion transformers (DiTs) for text-to-image synthesis.

Technical level: Advanced. The paper assumes familiarity with diffusion models, the DiT architecture, Mixture-of-Experts (MoE) routing, knowledge distillation, and pruning-based compression.

Scope: The paper proposes Dense2MoE, a framework that converts a dense DiT text-to-image model (FLUX.1 [dev], 12B parameters) into a sparse MoE model with fewer activated parameters while preserving generation quality.

What This Paper Is About

Large diffusion transformers such as FLUX.1 [dev] have 12 billion parameters — 13.8 times larger than SD1.5 — which makes inference memory-hungry and slow. Existing compression approaches rely mainly on pruning, which permanently removes parameters and therefore shrinks model capacity, causing severe quality loss at high compression rates. The paper's goal is to reduce the number of activated parameters at inference without reducing the model's total capacity, by dynamically routing each input through different subsets of parameters.

Key Contributions

  1. First dense-to-MoE conversion for diffusion models. The authors state this is the first attempt to apply a dense-to-MoE paradigm to diffusion models, reducing activated parameters while preserving model capacity.
  2. A unified structured sparsification framework. The work combines Mixture of Experts (MoE) at the FFN level with Mixture of Blocks (MoB) at the transformer-block level, so sparsity is achieved in both width and depth.
  3. A specialized knowledge distillation pipeline. The pipeline consists of Taylor-metric-based expert initialization, distillation with a load-balancing loss, and a group feature loss for MoB optimization.
  4. FLUX.1-MoE model family. The authors present four sparse models (L, M, S, XS) distilled from FLUX.1 [dev] with 5.2B, 4B, 3.2B and 2.6B activated parameters, reporting performance competitive with or better than pruning-based baselines of the same scale.

Main Findings

  • FFNs dominate parameter count. In DiT, FFNs account for nearly 50% of total parameters, which the authors identify as the main target for MoE conversion. Replacing the FFNs with MoE layers reduces the activated parameters in the FFNs by 62.5%.
  • Block importance is input-dependent. Using MSE between the input and output of each single-stream block in FLUX.1 [dev], the authors show that a block's contribution varies significantly across timesteps and prompts, which motivates sample-dependent block selection rather than average-based depth pruning.
  • Overall compression. The abstract reports a 60% reduction in activated parameters while maintaining original performance; the conclusion states compression of the 12B FLUX.1 [dev] to 5.2B activated parameters while maintaining original performance; the contributions section describes a reduction of over 56% in activation parameters.
  • Headline benchmark numbers (28 NFEs, 1024×1024). FLUX.1 [dev] baseline: 66.00 TFLOPs, 11.90B activated parameters, 21.20 s latency, CLIP 32.24, GenEval 0.6595, DPG 83.42. FLUX.1-MoE-L: 43.42 TFLOPs, 5.15B activated parameters, 17.80 s latency, CLIP 31.39, GenEval 0.5702, DPG 81.63. FLUX.1-MoE-XS: 20.26 TFLOPs, 2.64B activated parameters, 8.74 s latency, CLIP 30.40, GenEval 0.4036, DPG 73.66.
  • Comparison against FLUX.1-Lite. FLUX.1-MoE-L achieves better performance on multiple benchmarks with 3B fewer activated parameters and 20% fewer FLOPs (43.42 T vs 53.15 T for FLUX.1-Lite, which has 8.16B activated parameters).
  • Comparison at high compression. At a 75% compression rate of activated parameters, FLUX.1-MoE-S (3.19B activated, CLIP 30.67, GenEval 0.4441) and FLUX.1-MoE-XS (2.64B activated, CLIP 30.40, GenEval 0.4036) outperform FLUX-Mini (3.18B activated, 50 NFEs, CLIP 29.94, GenEval 0.3209).
  • Step-accelerated variants. HyperFLUX-MoE-L (8 NFEs) reaches 5.09 s latency at 5.15B activated parameters with CLIP 31.50, and HyperFLUX-MoE-XS (8 NFEs) reaches 2.50 s at 2.64B activated parameters with CLIP 30.70.
  • MoE beats MLP pruning at equal activation. Reducing the activated expansion ratio from 4 to 1.5 (62.5% compression), the MoE method scores CLIP 31.81, IR 0.8368, MPS 12.74, GenEval 0.5728, DPG 81.24, versus Diff-Pruning at r_a=1.5 (CLIP 30.89, IR 0.4456, MPS 11.56, GenEval 0.4113, DPG 72.23) and even versus Diff-Pruning at r_a=2.0 (CLIP 31.31, IR 0.6033, MPS 12.01, GenEval 0.4888, DPG 77.53).
  • MoB beats depth pruning at equal depth. With 9 double-stream and 38 single-stream blocks activated, MoB scores CLIP 31.59, IR 0.7903, MPS 12.72, GenEval 0.5396, DPG 78.86 versus Lite (31.48, 0.7491, 12.69, 0.4919, 77.72) and BK (31.20, 0.6849, 12.46, 0.4789, 76.39). With 9 double-stream and 26 single-stream blocks, the gap widens sharply: MoB (31.05, 0.6541, 12.27, 0.4956, 76.51) versus Lite (26.64, −1.1657, 7.79, 0.0926, 41.62) and BK (30.31, 0.2587, 11.13, 0.3450, 66.86).
  • Expert configuration matters. The configuration with a larger shared expert (r_s=1, r_n=0.25, n=12, k=2) outperforms the one with a smaller shared expert (r_s=0.5, r_n=0.25, n=14, k=4) at both the shared-only stage (CLIP 31.44 vs 31.32) and the full MoE stage (CLIP 31.52 vs 31.48).
  • Taylor initialization and staged distillation help. Replacing the Taylor metric with random weight splitting degrades shared-expert results (−0.02 CLIP, −0.159 IR, −0.17 MPS, −0.0013 GenEval), and distilling shared and normal experts jointly instead of separately also degrades results (−0.04 CLIP, −0.0204 IR, −0.07 MPS, −0.0057 GenEval).
  • More MoB groups work better. Among three settings that each activate 4 of 12 blocks, more groups (4 groups of 3 blocks) gives better distillation performance than 2 groups of 6 or a single group of 12, attributed to more supervision layers from the group feature loss.
  • Experts show interpretable specialization. Sampling with 1K MJHQ-30K prompts organized into 10 categories, the authors observe that in the double-stream block's image branch, expert selection patterns align with prompt categories, shift across timesteps (more concentrated in the high-noise stage, especially step 0), and follow a raster-like spatial structure across tokens. In the single-stream block, most text tokens are empty due to prompt length, producing convergent expert selection.
  • Dynamic Top-K works without retraining. Because of the distillation pipeline, the model supports changing the number of activated normal experts at inference. With zero activated normal experts the output loses detail and shows higher color saturation; increasing the count improves detail and realism.

Methodology in Plain English

The authors take a trained dense DiT and rebuild it as a sparse MoE of the same total size, then distill the original model's behavior back into the sparse one.

Step one: replace FFNs with MoE layers. Each feed-forward network is replaced by a shared expert plus multiple smaller normal experts controlled by a gating network. During inference only the top-k normal experts plus the shared expert are activated per token, so the activated expansion ratio r_a = r_s + k · r_n is smaller than the dense ratio r = r_s + n · r_n.

Step two: initialize experts smartly. Rather than creating experts randomly, the authors rank MLP weights by a first-order Taylor importance score, slice the MLP into segments along the intermediate feature dimension, place the highest-scoring segments into the shared expert, and distribute the rest evenly among normal experts. The shared expert is then improved by knowledge distillation, treating it temporarily as a pruned model.

Step three: distill the full MoE. The normal experts and the gating network are activated, shared experts are frozen, and training uses an output distillation loss, a per-layer block feature loss with normalized weights, and a load-balancing loss (weighted 10^-2) to prevent the router from collapsing onto a few experts.

Step four: group blocks into a Mixture of Blocks. Consecutive transformer blocks are grouped, and a block router selects only some blocks within each group to run. The router reuses the global AdaLN conditioning embedding along with the token features to decide, so selection can depend on text and timestep. A group feature loss aligns the last block of each original group with the output of the compressed group, isolated blocks are frozen, and load balancing is applied.

Why This Matters

Research impact. The paper reframes diffusion model compression as a capacity-preserving problem rather than a capacity-reducing one. By keeping total parameters constant and varying which ones are active, it offers an alternative to pruning that the authors show outperforms pruning baselines at matched activation levels.

Real-world applications:

  • On-device or consumer-GPU text-to-image generation, where reducing activated parameters and latency (e.g., 8.74 s at 2.64B activated parameters versus 21.20 s at 11.90B) directly lowers hardware requirements.
  • Cost-efficient cloud inference for image generation services, where FLOPs and latency reductions translate into lower serving cost.
  • Content creation pipelines that need fast iteration, especially with the 8-NFE step-distilled variants reaching 2.50 s latency.
  • Flexible quality/latency trade-offs at deployment time, since the same model supports different Top-K activation settings without retraining.

Industry relevance. The work comes from a collaboration between Sun Yat-sen University and ByteDance (Intelligent Creation and Seed Vision), and it targets a production-scale model, FLUX.1 [dev]. The MoE paradigm is already dominant in large language models, so demonstrating it for a large DiT connects generative image infrastructure to the same efficiency playbook used for LLMs.

Future Directions

  • Pushing compression further. The authors frame additional sparsity as an open avenue, noting that the inherent flexibility of MoE "opens up new avenues for further exploration."
  • Understanding and exploiting dynamic Top-K. Since the model retains basic generative ability with zero normal experts activated, systematic study of quality/latency trade-offs at inference time without retraining remains open.
  • Improving very high depth compression. The degradation of the single-block distillation baselines (e.g., Lite dropping to IR −1.1657 and DPG 41.62 when further reducing single-stream blocks) suggests that feature-alignment strategies for aggressive depth removal are not yet solved.
  • Better MoB routing. The finding that more MoB groups improve results, and that routing currently reuses the AdaLN embedding, leaves room for improved routing mechanisms tailored to text-to-image denoising.

Target Audience

Researchers and engineers working on efficient diffusion models, diffusion transformer architectures, and model compression; practitioners deploying large text-to-image models who need to reduce inference cost; and readers interested in Mixture-of-Experts methods beyond language models, including MoE interpretability and routing behavior.

Authors’ abstract

Diffusion Transformer (DiT) has demonstrated remarkable performance in text-to-image generation; however, its large parameter size results in substantial inference overhead. Existing parameter compression methods primarily focus on pruning, but aggressive pruning often leads to severe performance degradation due to reduced model capacity. To address this limitation, we pioneer the transformation of a dense DiT into a Mixture of Experts (MoE) for structured sparsification, reducing the number of activated parameters while preserving model capacity. Specifically, we replace the Feed-Forward Networks (FFNs) in DiT Blocks with MoE layers, reducing the number of activated parameters in the FFNs by 62.5\%. Furthermore, we propose the Mixture of Blocks (MoB) to selectively activate DiT blocks, thereby further enhancing sparsity. To ensure an effective dense-to-MoE conversion, we design a multi-step distillation pipeline, incorporating Taylor metric-based expert initialization, knowledge distillation with load balancing, and group feature loss for MoB optimization. We transform large diffusion transformers (e.g., FLUX.1 [dev]) into an MoE structure, reducing activated parameters by 60\% while maintaining original performance and surpassing pruning-based approaches in extensive experiments. Overall, Dense2MoE establishes a new paradigm for efficient text-to-image generation.

Read the original paper