Skip to content
AI.info

Research

H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion Models

H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion Models Overview Research area: Computer Vision — inference acceleration for generative diffusi

H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion Models
arXiv
2510.27171
Published
2025-10-31
Authors
Mingyu Sung, Il-Min Kim, Sangseok Yun, Jae-Mo Kang

AI summary

H2-Cache: A Novel Hierarchical Dual-Stage Cache for High-Performance Acceleration of Generative Diffusion Models

Overview

Research area: Computer Vision — inference acceleration for generative diffusion models via caching mechanisms.

Technical level: Intermediate (assumes familiarity with diffusion sampling, transformer block structure, and similarity-threshold caching).

Scope: The paper proposes and empirically evaluates H2-Cache, a two-stage, dual-threshold caching method for the Flux diffusion architecture, reporting a 5.08x speedup at near-baseline image quality on the CUTE80 dataset.

What This Paper Is About

Diffusion models generate images through a long, iterative denoising loop, which makes inference slow and expensive. Existing caching tricks that skip entire network blocks can speed this up, but they tend to blur fine details and add so much per-block checking overhead that the speedup can disappear. The authors' goal is to design a caching scheme that skips redundant computation selectively — protecting detail-rich computations while still cutting a large share of the work — so that generation gets much faster without visibly degrading the image.

Key Contributions

  1. A functional dichotomy within the denoising network. The authors identify and exploit a separation between a structure-defining stage (B_L1, the multi-transformer block processing that produces an intermediate feature map z'_t) and a detail-refining stage (B_L2, the single-transformer block processing that produces the noise prediction).

  2. H2-Cache itself, a hierarchical two-stage caching mechanism that applies independent thresholds (τ1 for the structural check, τ2 for the detail check) to each stage, rather than treating the block as a monolithic unit.

  3. Pooled Feature Summarization (PFS), a lightweight technique for fast similarity estimation between high-dimensional tensors. PFS downsamples each tensor with hardware-accelerated average pooling into a "thumbnail" and computes a relative difference metric on the compact representations, making the doubled frequency of cache checks computationally feasible.

  4. Extensive experiments on the modern Flux architecture showing up to 5.08x acceleration with near-baseline quality, outperforming block cache and TeaCache on quantitative metrics such as CLIP-IQA and on perceptual quality.

Main Findings

  • Aggregate speed and quality trade-off (Table 1, CUTE80 dataset): Baseline averaged 55.72 s (1.0x) with CLIP-IQA 0.7693; Block Cache 12.82 s (4.35x) with 0.7681 (-0.16%); teaCache 11.16 s (4.99x) with 0.7462 (-3.00%); H2-Cache 10.97 s (5.08x) with 0.7688 (-0.07%). Thresholds were set to τ = 1.0 for both Block Cache and teaCache.

  • Per-prompt comparison (Fig. 3): The no-caching baseline required over 53 seconds, while all caching methods reduced inference to under 12 seconds. For the first prompt, Block Cache and TeaCache dropped CLIP-IQA from the baseline's 0.9800 to 0.9448 and 0.9116 respectively, while H2-Cache reached the fastest time of 9.42 s with a CLIP-IQA of 0.9707. For the second prompt, H2-Cache achieved a 7.33x speedup (54.19 to 7.39 s) with quality moving from 0.9868 to 0.9346.

  • Functional separation confirmed qualitatively (Fig. 2): Caching B_L1 freezes the overall pose and layout, whereas caching B_L2 preserves fine-grained textures while allowing the global structure to evolve — empirically validating the two-stage decomposition.

  • Hierarchical thresholds interact (Fig. 4 ablation): A low structural threshold of τ1 = 0.15 achieved the best peak performance across metrics — PSNR 17.30 and SSIM 0.77 at τ2 = 0.17, CLIP-IQA 0.77 at τ2 = 0.16, and FID 60.22 (lower is better) at τ2 = 0.19 — but these peaks are sharp and sensitive to τ2. Higher τ1 values gave more stable but sub-optimal performance, so there is no single universally optimal setting.

  • PFS delivers speedup with marginal quality cost (Table 2): With τ1 = 0.15 fixed, adding PFS reduced processing time by up to 14.5%, while changes in PSNR and SSIM generally remained below 3%.

  • PFS stabilizes the difference signal (Fig. 5): Without PFS the difference metric fluctuated significantly; with PFS (e.g., D_p1 = 512, D_p2 = 384) the metric was visibly more stable across inference steps, reducing generation time from 12.11 to 11.15 s in one case and from 10.22 to 9.78 s in another.

  • Speedup grows with step count (Table 3): With τ1 = 0.15 and τ2 = 0.18, the speedup over a no-cache baseline scaled from 1.22x at 10 steps (4.89 s, CLIP-IQA 0.45, -13.74%) to 5.08x at 100 steps (10.97 s, CLIP-IQA 0.77, -0.07%), indicating that computational savings increasingly outweigh the cache-hit detection overhead at larger step sizes.

Methodology in Plain English

The authors start from an observation about how the Flux architecture works: a single denoising step is naturally split into a first chunk of transformer blocks that establishes the image's overall composition and a final block that fills in fine detail. Instead of caching or skipping the whole chunk as one unit, they keep two separate checks.

At each timestep, they first compare the current noisy latent against the latent from the last fully computed step. If the difference (measured by L2 distance) is below the primary threshold τ1, they conclude the structure has stabilized and reuse both the intermediate features and the final noise prediction, skipping all computation for that step. If that check fails, they compute the structural stage anyway, but then run a second check: they compare the freshly computed intermediate features against the cached ones. If that difference is below the secondary threshold τ2, they reuse the cached noise prediction and skip the detail-refining block; otherwise they compute it. Whenever any part of the block is recomputed, the whole cache tuple is refreshed.

Because this means running similarity checks twice per step, a naive full-tensor comparison would defeat the purpose. Their fix — PFS — shrinks each tensor with average pooling into a small "thumbnail" before comparing, using a pooling kernel derived from the tensor height divided by a per-stage divisor (D_p1 for the structural stage, D_p2 for the detail stage). The final decision uses a relative difference metric between the thumbnails. Pooling is hardware-accelerated, and averaging over local patches makes the signal less sensitive to high-frequency noise, so the cache decisions are both cheap and stable.

Experiments used the nunchaku framework, the flux.1-dev quantized model (nunchaku-flux.1-dev/svdq-int4_r32-flux.1-dev), 1024x1024 resolution, guidance scale 3.5, batch size 1, 100 inference steps, a single NVIDIA A5000 GPU, and an Intel Core i9-14900K CPU. Prompts came from the CUTE80 dataset via LLaVA-NeXT-8B, with the XLabs-AI/flux-RealismLora weights integrated at LoRA strength 0.8.

Why This Matters

Impact on research: The work reframes caching for diffusion models as a structured problem rather than a uniform thresholding problem, showing that knowing what a block computes changes how it can be safely skipped. It also offers a practical alternative to learned caching policies like TeaCache's teacher network, avoiding the overhead of training a separate policy network.

Real-world applications:

  • Real-time or interactive image generation tools, where multi-second latency currently blocks responsive use.
  • Creative and design workflows where users iterate on prompts and need fast turnaround at high fidelity.
  • Content pipelines that must generate large volumes of high-resolution imagery on limited GPU fleets.
  • On-device or cost-constrained deployment, where the reported 5.08x reduction in compute directly lowers hardware requirements.

Industry relevance: Serving diffusion models is GPU-bound, so a roughly 5x reduction in inference time at near-identical perceptual quality is a direct cost and throughput win for any product built on text-to-image generation. The fact that the method is demonstrated on a quantized Flux variant on a single NVIDIA A5000 makes the results directly relevant to commodity deployment rather than only large-scale clusters.

Future Directions

  • Learnable thresholds. The current implementation relies on empirically set τ1 and τ2; the authors propose developing a learnable policy to automate their selection for better adaptability.
  • Extension to other modalities. Applying the core principles of H2-Cache to accelerate video and 3D diffusion models is flagged as an open avenue.
  • Robustness of the threshold trade-off. The ablation shows top performance is sharply sensitive to τ2 when τ1 is low, raising the question of how to keep peak performance without that fragility across prompts.
  • Low-step regimes. Speedup drops to 1.22x at 10 steps (with a -13.74% CLIP-IQA change), so the method is not yet advantageous when few denoising steps are used — an unresolved limitation.

Target Audience

Researchers and engineers working on diffusion model inference efficiency, model serving, and generative AI deployment. The paper is also useful for practitioners applying caching to transformer-based generative architectures such as Flux, and for students with background in diffusion sampling and latent diffusion models who want a concrete example of architecture-aware optimization.

Authors’ abstract

Diffusion models have emerged as state-of-the-art in image generation, but their practical deployment is hindered by the significant computational cost of their iterative denoising process. While existing caching techniques can accelerate inference, they often create a challenging trade-off between speed and fidelity, suffering from quality degradation and high computational overhead. To address these limitations, we introduce H2-Cache, a novel hierarchical caching mechanism designed for modern generative diffusion model architectures. Our method is founded on the key insight that the denoising process can be functionally separated into a structure-defining stage and a detail-refining stage. H2-cache leverages this by employing a dual-threshold system, using independent thresholds to selectively cache each stage. To ensure the efficiency of our dual-check approach, we introduce pooled feature summarization (PFS), a lightweight technique for robust and fast similarity estimation. Extensive experiments on the Flux architecture demonstrate that H2-cache achieves significant acceleration (up to 5.08x) while maintaining image quality nearly identical to the baseline, quantitatively and qualitatively outperforming existing caching methods. Our work presents a robust and practical solution that effectively resolves the speed-quality dilemma, significantly lowering the barrier for the real-world application of high-fidelity diffusion models. Source code is available at https://github.com/Bluear7878/H2-cache-A-Hierarchical-Dual-Stage-Cache.

Read the original paper