Skip to content
AI.info

Research

RSTR: Reducing SpatioTemporal Redundancy in Diffusion Transformers

Overview Research area: Efficient inference for diffusion transformers (DiT) in image generation — specifically, eliminating redundant Classifier-Free Guidance (CFG) computation and redundant feature-

RSTR: Reducing SpatioTemporal Redundancy in Diffusion Transformers
arXiv
2512.14096
Published
2025-12-16
Authors
Ruitong Sun, Tianze Yang, Wei Niu, Jin Sun

AI summary

Overview

  • Research area: Efficient inference for diffusion transformers (DiT) in image generation — specifically, eliminating redundant Classifier-Free Guidance (CFG) computation and redundant feature-calibration computation.
  • Technical level: Advanced. The paper assumes familiarity with diffusion sampling, classifier-free guidance, transformer block caching, SVD-based low-rank approximation, and evolutionary/black-box optimization.
  • Scope: A two-stage framework (RSTR) that jointly optimizes when to apply CFG and what guidance scale to use, then allocates non-uniform caching calibration ranks across transformer regions, evaluated on DiT-XL/2, PixArt-α, FLUX, and Qwen-Image.

What This Paper Is About

Diffusion transformers produce high-quality images but are expensive: iterative denoising costs trillions of floating-point operations, and Classifier-Free Guidance doubles that cost by running a conditional and an unconditional forward pass at every single timestep. The authors argue this cost hides two distinct kinds of waste — temporal redundancy (guidance is applied uniformly across timesteps even though it only matters at a few) and spatial redundancy (caching-based calibration is applied with the same rank to every transformer block, even though blocks differ in sensitivity). The goal is to remove both simultaneously, since existing efficiency methods sacrifice quality while existing quality methods save no computation.

Key Contributions

  1. Temporal optimization: The first framework to jointly optimize discrete CFG skip patterns and continuous guidance scales via evolutionary search, reducing CFG evaluations by up to 80% while maintaining generation quality.
  2. Spatial optimization: An adaptive rank allocation scheme that assigns different calibration capacities to different transformer regions based on their sensitivity, enabling effective caching under variable guidance where uniform-rank methods fail.
  3. Insights into guidance dynamics: An analysis showing that sparse high-scale CFG can effectively replace dense low-scale CFG, but that guidance position and scale are inherently coupled and every high-scale step is indispensable — explaining why learned search outperforms analytical derivation.
  4. Empirical validation across four model families: Results on DiT-XL/2, PixArt-α, FLUX, and Qwen-Image showing 50%–70% compute savings while maintaining or improving quality.

Main Findings

  • DiT-XL/2 at 512×512: RSTR applies guidance at only 9 of 50 timesteps, reaching FID 2.72 versus 3.20 for 50-step DDIM, with 47% less computation reported (24.97T versus 52.45T MACs), and surpassing 1000-step DDIM (FID 2.99). Reported MACs drop from 52.5 (DDIM 50) to 24.9 with adaptive caching, and latency from 20.86 to 11.28.
  • DiT-XL/2 at 256×256: RSTR achieves the best FID among reported baselines (2.04), ahead of 1000-step DDIM (2.12), with 8 of 50 timesteps using CFG and 5.1 MACs.
  • PixArt-α on MSCOCO (256×256): Only 6 of 20 timesteps require guidance, giving FID 19.27 (a 21.7% improvement) with 60% computational reduction (2.67 versus 6.72 MACs) and preserved text-image alignment (CLIP 16.48 versus 16.31).
  • FLUX at 512×512 with True CFG: Only 8 of 20 timesteps require guidance. Without caching, RSTR achieves the highest ImageReward (1.0092) and CLIP Score (28.06) while using 72% less computation than the 50-step full-CFG baseline; GenEval reaches 67.46 versus 68.60 for full CFG.
  • Qwen-Image at 1024×1024: Only 10 of 20 timesteps require guidance, reported as a 3.43× latency speedup in the abstract and a 4.7× speedup (56.73s to 16.62s) in the results section, with GenEval 87.21 versus 88.40 and improved GenEval2 GM (40.57 versus 38.72).
  • Guidance concentration principle: Within a 12-step interval (steps 12–23) of the DiT-XL/2 256×256 schedule, dense low-scale guidance (w=1.4, 12 evaluations) yields FID 2.10; reducing density while raising scale (w=1.8/6 evals, w=2.2/4 evals, w=7.4/1 eval) keeps FID essentially flat (2.11, 2.10, 2.10), confirming a 12× CFG reduction inside that interval.
  • Position and scale are coupled: Shifting the w=7.4 step from step 12 to step 16 improves IS (272.4) but worsens FID (2.21); shifting earlier to step 8 lowers IS (256.5) but improves FID (2.04). Removing any high-scale step degrades both metrics (e.g., removing step 39 gives IS 230.8, FID 2.44 versus 263.0 and 2.10 with none removed).
  • Graceful scaling behavior: Under a 3.3× multiplicative scaling of the schedule, constant CFG degrades severely (FID 2.23 to 16.39) while RSTR degrades far less (2.10 to 9.20), suggesting sparse schedules filter harmful guidance signals.
  • Cross-model-size transfer: A schedule optimized on DiT-XL/2 (675M) applied without re-optimization to DiT-B/2 (131M) and DiT-S/2 (33M) improves FID — 11.52 to 9.22 (−20%) and 24.09 to 20.00 (−17%) respectively — with consistent sFID and Precision gains.
  • Why uniform caching fails: The paper reports elevated feature reconstruction MSE across all transformer blocks under variable guidance, most severe in deeper blocks. Rank r=256 versus r=512 shows a counter-intuitive pattern: r=512 suppresses error spikes in Block 22 but produces higher errors in Blocks 0, 14, and 26; only r=1024 consistently minimizes error across all blocks, at prohibitive cost.
  • Reference trajectory length: Table 9 (partially shown for DiT-XL/2 512×512) reports RSTR† at T_ref=50 with 50 steps, 33.56 MACs, IS 229.6, FID 2.84, sFID 4.40, Precision 83.80, and Recall 56.0, compared with DDIM 1000 (1049.1 MACs, IS 210.6, FID 2.99) and DDIM 50 (52.45 MACs, IS 203.8, FID 3.20).

Methodology in Plain English

The approach is split into two sequential stages.

Stage 1 — deciding when to guide and how strongly. The constant guidance scale is replaced by a per-timestep schedule, where each timestep's scale sits between 1 and a maximum. If a timestep's scale falls below a threshold, the unconditional forward pass is dropped entirely and only the conditional pass runs. Finding the best schedule is a hybrid discrete-continuous problem: gradient methods are impractical because backpropagation would have to traverse the entire denoising trajectory. Instead, the authors run an evolutionary search in a transformed (sigmoid) space for numerical stability. A population center defines base guidance values; candidates are sampled by adding Gaussian noise whose scale shrinks over generations; each candidate is scored by a fitness combining output-matching quality against a longer reference trajectory (typically 100–1000 steps versus 20–50) and a sparsity reward. The center is updated using rank-based weighted averaging, in the style of established evolution strategies.

Stage 2 — deciding how much caching calibration each region needs. Sparse schedules create larger feature differences between consecutive timesteps, which the standard incremental-calibration caching equation cannot absorb. The paper derives two specific deviation sources: consecutive CFG steps with different scales, and transitions between CFG and conditional-only steps. Because these errors are unevenly distributed across blocks, the transformer is partitioned into K equal regions, each getting its own SVD truncation rank for the calibration matrix. Ranks are chosen by coordinate descent — fix all other regions, tune one region's rank by binary search to minimize FID, subject to a total rank budget — and iterated until the configuration converges.

Why This Matters

The work shows that guidance scheduling and feature caching are not independent optimizations that can simply be stacked; the first creates the conditions that break the second, and the second must adapt. This reframes diffusion acceleration as a coupled spatiotemporal problem rather than two separate efficiency tricks, and provides a concrete mechanism (adaptive, sensitivity-based rank allocation) that uniform-rank caching methods lack.

  • Interactive image generation: Reduced latency (reported 3.43× on Qwen-Image at 1024×1024, with 56.73s to 16.62s reported for one comparison) directly affects responsiveness of text-to-image tools used by designers and consumers.
  • Content creation and advertising pipelines: FLUX results on DrawBench and GenEval indicate that higher ImageReward and CLIP alignment can be retained at a fraction of the compute, relevant for batch generation of marketing or editorial imagery.
  • On-device and edge deployment: Because the method requires no architectural modification and keeps full denoising steps, it is compatible with existing diffusion transformer checkpoints, which matters for resource-constrained hardware.
  • Compute-budget-constrained research and medical imaging: The paper cites prior work on medical and health-related diffusion, where fewer MACs per image makes larger-scale generation studies feasible.
  • Industry relevance: The compute savings translate into GPU-hour and energy cost reductions for cloud inference providers, and the schedule-transfer results (across FLUX and Qwen-Image in the appendix, and across DiT sizes in the main text) reduce the cost of re-tuning for each new model.

Future Directions

  • Automating schedule discovery across architectures: The paper notes that optimal guidance schedules vary across architectures, and demonstrates transfer within the DiT family; whether a single search procedure can produce portable schedules across model families remains open.
  • Generalizing adaptive rank allocation: The current scheme partitions blocks into K equal regions by network position; whether non-uniform, learned, or sensitivity-clustered partitions outperform uniform division is not addressed.
  • Replacing search with analysis: The authors state that position and scale are coupled and non-linear, which is why learned optimization beats analytical derivation; deriving a closed-form characterization of the guidance-scale/coverage relationship would eliminate the search cost.
  • Extending temporal skipping beyond CFG: The paper targets unconditional passes and per-step caching; whether the same joint formulation applies to step-count reduction, distillation, or quantization settings is not reported.

Target Audience

Researchers and engineers working on diffusion model inference efficiency, generative model serving infrastructure, and transformer acceleration. It is most useful to readers already comfortable with CFG, diffusion samplers, and caching-based acceleration, since the contributions are framed as fixes to specific limitations of those existing techniques. Readers seeking introductory material on diffusion models will find the paper assumes substantial background.

Authors’ abstract

Diffusion Transformers (DiTs) have achieved remarkable success in image generation, yet their deployment is hindered by high computational costs. We identify two sources of redundancy. First, temporal redundancy: Classifier-Free Guidance (CFG) applies costly dual forward passes at every timestep, yet guidance matters only at specific steps, and variable scales at critical steps can compensate for skipping others. Second, spatial redundancy: under variable guidance, different transformer blocks exhibit heterogeneous sensitivity, yet uniform calibration across all blocks wastes computation while failing to address their varying requirements. We present RSTR, the first framework to jointly reduce spatiotemporal redundancy in diffusion transformers. Stage-1 addresses temporal redundancy through evolutionary search, discovering sparse guidance schedules with variable scales. Stage-2 addresses spatial redundancy through adaptive rank allocation, assigning calibration capacities to transformer regions based on their sensitivity. Experiments on DiT-XL/2, PixArt-$α$, FLUX, and state-of-the-art Qwen-Image demonstrate 50%-70% compute savings while maintaining or improving quality. On DiT-XL/2, RSTR achieves 57% savings with 15% FID improvement; on Qwen-Image, 3.43$\times$ speedup with preserved quality.

Read the original paper