Skip to content
AI.info

Research

ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer Acceleration

ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer Acceleration Overview Research area: Efficient generative modeling — specifically training-free inferenc

ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer Acceleration
arXiv
2512.17298
Published
2025-12-19
Authors
Fanpu Cao, Yaofo Chen, Zeng You, Wei Luo

AI summary

ProCache: Constraint-Aware Feature Caching with Selective Computation for Diffusion Transformer Acceleration

Overview

  • Research area: Efficient generative modeling — specifically training-free inference acceleration for Diffusion Transformers (DiTs) used in image and video synthesis.
  • Technical level: Advanced. The paper assumes familiarity with diffusion sampling, transformer block internals (self-attention, cross-attention, MLP, AdaLN), and caching-based inference acceleration.
  • Scope: A training-free dynamic feature caching framework, ProCache, that combines an offline constraint-aware caching-pattern search with a selective (partial layer and token) recomputation module, evaluated on PixArt-α, DiT-XL/2, and the FLUX.1 family.

What This Paper Is About

Diffusion Transformers produce state-of-the-art generative results but are computationally expensive because they run many sequential denoising steps. Feature caching speeds them up by reusing intermediate features across steps instead of recomputing them, but existing methods reuse features at fixed uniform intervals and without any correction, which causes error accumulation and quality loss. ProCache instead searches for a model-tailored, non-uniform caching schedule and selectively recomputes only a small subset of deep layers and high-importance tokens during cached steps, achieving up to 1.96× and 2.90× acceleration on PixArt-α and DiT respectively with negligible quality degradation.

Key Contributions

  1. An analysis showing that DiT feature evolution is non-uniform across denoising steps (stable early and mid stages, highly dynamic in later steps) and that error propagation is concentrated in deeper blocks, revealing a mismatch with uniform caching schedules.
  2. ProCache, a training-free acceleration framework built from two components: a constraint-aware caching pattern search and a selective computation strategy.
  3. A constraint-aware caching pattern search module that generates non-uniform activation schedules through offline constrained sampling under three constraints (budget, monotonic, bounded), requiring less than one hour on a single GPU and no training.
  4. A selective computation module that refreshes only the deepest blocks and the top-p% most important tokens during cached segments, adding only about 3% latency overhead.

Main Findings

  • DiT error grows with depth: Relative L1 error across diffusion steps in DiT blocks (computed from 10 samples on PixArt-α) grows progressively, with deeper blocks (e.g., Block 25–28) showing significantly higher magnitudes than shallower ones (e.g., Block 1–4).
  • Feature divergence is temporally skewed: Output L1 error between the current step and the previous step in DiT-XL/2 stays small in early and mid stages but grows sharply in later steps, following an approximately exponential trend.
  • ImageNet (DiT-XL/2) results: ProCache achieves 2.90× speedup with FID 2.96, sFID 4.93, Precision 0.80, Recall 0.57, Inception Score 232.85, latency 1.725 s and 8.18 T FLOPs. For comparison, FORA (𝒩=3) reaches 2.76× with FID 3.88 and sFID 6.43; ToCa (𝒩=3) reaches 2.32× with FID 3.04; ToCa (𝒩=4) reaches 2.72× with FID 3.64; Δ-DiT (𝒩=3) reaches 1.47× with FID 3.75.
  • FID/sFID gain claimed over uniform caching: The paper states ProCache achieves nearly a 30% improvement in FID and sFID over uniform caching (FORA) while enhancing the speedup ratio by 5%.
  • PixArt-α on MS-COCO2017: ProCache reaches 1.96× speedup, latency 1.215 s, 5.70 T FLOPs, FID 27.66, and CLIP 16.45. FORA² reaches 2.79× at FID 29.84, and Δ-DiT reaches 1.54× at FID 28.91.
  • FLUX.1-dev on PartiPrompts: ProCache reaches 1.54× speedup (18.73 s, 2415.25 T) with Image Reward 1.207, versus FORA at 1.51× with 1.196 and ToCa (𝒩=2) at 1.51× with 1.202.
  • FLUX.1-schnell on PartiPrompts: ProCache (𝒩=2) reaches 1.56× speedup (1.817 s, 177.26 T) with Image Reward 1.138, versus ToCa (𝒩=2) at 1.53× with 1.134. Pattern search was not used here because the model generates in just four steps; a fixed interval of 2 steps was used instead.
  • Component ablation (DiT-XL/2): From a default configuration (B=13) at 2.93× with FID 4.75 and sFID 8.43, adding the searched pattern keeps 2.93× but improves to FID 3.15 and sFID 5.12; adding selective computation alone gives 2.90×, FID 3.28, sFID 5.95; the full ProCache gives 2.90×, FID 2.94, sFID 4.93.
  • Sampling budget ablation: As K increases from 5 to 15, FID improves from 45.07 to 45.02 to 44.96 with Inception Scores 181.69, 182.50, and 181.63 respectively; K=5 is adopted as the default because even the smallest budget surpasses existing approaches.
  • Block-ratio ablation: With the proportion r of computed blocks at 50%, 65%, 75%, and 90%, FLOPs decrease (9.138, 8.911, 8.594, 8.344 T) while FID stays near 45.3 (45.32, 45.33, 45.31, 45.38) and Inception Score falls from 184.18 to 182.46; 75% of blocks is reported as achieving the lowest FID.
  • Robustness: Across five run trials on ImageNet-1K, variation in FID and Inception Score across runs is minimal, indicating the search reliably finds promising caching patterns.
  • High-acceleration regime: On ImageNet-50k, ProCache is reported to reduce quality degradation by 56.2% and to perform robustly at acceleration ratios exceeding 4.53×.

Methodology in Plain English

The authors start by measuring how much DiT internal features change over the denoising process. They find two things: errors pile up unevenly across the network, concentrating in the deeper blocks, and features change slowly at first and then rapidly near the end of sampling. Uniform caching (recompute every N steps) therefore wastes compute early and loses accuracy late.

ProCache has two parts. First, instead of a fixed interval, it represents the whole inference as a binary string where 1 means "compute this step" and 0 means "reuse the cache." It randomly samples many such strings offline but keeps only those satisfying three constraints: a total computation budget B, reuse intervals that never grow over time (monotonic), and intervals falling inside a minimum/maximum range (bounded). Each surviving candidate is scored on a small dataset with a quality metric such as FID, and the best one becomes the schedule. This costs less than an hour on a single GPU and involves no training.

Second, within cached stretches, the method does not blindly reuse everything. At every second position inside each contiguous zero block, it recomputes only a subset of layers — the deepest ones, where errors concentrate — and only the top-p% tokens by attention-output L2 norm, since tokens with larger feature changes most affect the output. Self-attention is exempted from token selection because each token attends to all others, making partial computation prone to error propagation. Everything else is read from cache. Across the tested models this selective path adds roughly 3% latency while operating over a small number of caching steps, a subset of layers (25%), and 7–30% of tokens.

Why This Matters

  • Impact on research: The paper reframes feature caching as a scheduling and error-correction problem rather than a fixed-interval heuristic, and provides a training-free route to higher acceleration ratios than uniform caching baselines such as FORA, Δ-DiT, and ToCa.
  • Real-world applications:
    • Real-time or interactive text-to-image generation, where per-image latency directly determines usability.
    • High-resolution generation pipelines (the paper evaluates 1024×1024 outputs with FLUX) for design and media workflows.
    • Class-conditional and content-creation tools that need to synthesize large image batches (the paper generates 50,000 ImageNet images) under compute budgets.
    • Video synthesis, which the paper lists among diffusion model applications and which faces the same temporal redundancy and cost pressures.
  • Industry relevance: The method is plug-and-play and training-free, so it can be layered onto existing pretrained DiT checkpoints without fine-tuning and without retraining infrastructure. FLOPs and latency reductions reported in the paper (e.g., 8.18 T and 1.725 s on DiT-XL/2 versus 23.74 T and 4.549 s for DDIM-50) map directly to serving cost. The work was supported by the Postdoctoral Fellowship Program of CPSF (GZC20251043) and the Young Scholar Project of Pazhou Lab (No. PZL2021KF0021), and code is released at https://github.com/macovaseas/ProCache.

Future Directions

  • Whether the offline pattern search transfers across model checkpoints and tasks without re-searching, given that each model in the paper required its own hyperparameter configuration (Table A lists distinct values of K, B, p, v_min, v_max, and r for DiT-XL/2, PixArt-α, and FLUX.1-dev).
  • Extending the constraint-aware search to models with very few denoising steps: FLUX.1-schnell generates in four steps, and the paper abandoned pattern search entirely in favor of a fixed two-step interval, leaving the small-step regime unresolved.
  • Generalizing or scaling the selective computation design, for example the exemption of the self-attention module from token-level selection, which the authors justify by error-propagation risk but do not replace with an alternative.
  • Combining ProCache with orthogonal sampling-step reduction methods such as DPM-Solver, Rectified Flow, and DDIM, which the paper explicitly describes as orthogonal to feature caching but does not jointly evaluate.

Target Audience

Researchers and engineers working on diffusion model inference efficiency, particularly those deploying Diffusion Transformers under latency or compute constraints. It is also relevant to practitioners of training-free acceleration techniques (caching, token/layer skipping) and to readers interested in where error accumulates inside transformer-based diffusion networks. Beginners will need background in diffusion sampling and transformer architecture to follow the constraints, notation, and block-level analysis.

Authors’ abstract

Diffusion Transformers (DiTs) have achieved state-of-the-art performance in generative modeling, yet their high computational cost hinders real-time deployment. While feature caching offers a promising training-free acceleration solution by exploiting temporal redundancy, existing methods suffer from two key limitations: (1) uniform caching intervals fail to align with the non-uniform temporal dynamics of DiT, and (2) naive feature reuse with excessively large caching intervals can lead to severe error accumulation. In this work, we analyze the evolution of DiT features during denoising and reveal that both feature changes and error propagation are highly time- and depth-varying. Motivated by this, we propose ProCache, a training-free dynamic feature caching framework that addresses these issues via two core components: (i) a constraint-aware caching pattern search module that generates non-uniform activation schedules through offline constrained sampling, tailored to the model's temporal characteristics; and (ii) a selective computation module that selectively computes within deep blocks and high-importance tokens for cached segments to mitigate error accumulation with minimal overhead. Extensive experiments on PixArt-alpha and DiT demonstrate that ProCache achieves up to 1.96x and 2.90x acceleration with negligible quality degradation, significantly outperforming prior caching-based methods.

Read the original paper