Skip to content
AI.info

Research

GRACE: Generation-aware latent compression for efficient video generation

GRACE: Generation-Aware Latent Compression for Efficient Video Generation Overview Research area: Computer vision, specifically efficient video diffusion models, latent-space autoencoder compression,

GRACE: Generation-aware latent compression for efficient video generation
arXiv
2610.10524
Published
2026-10-07
Authors
Jiyoung Kim, Paul Hyunbin Cho, Jisu Nam, Donghoon Lee, Hyunsung Go, Yeonkyeong Lee, Hansaem Kim, Seungryong Kim

AI summary

GRACE: Generation-Aware Latent Compression for Efficient Video Generation

Overview

Research area: Computer vision, specifically efficient video diffusion models, latent-space autoencoder compression, and diffusion transformer (DiT) adaptation.

Technical level: Advanced. The paper builds on latent video diffusion, rectified flow / flow matching, causal 3D video autoencoders, LoRA adaptation, and feature-space alignment losses. Readers without background in diffusion models and latent compression will find the equations dense, though the high-level argument is accessible.

Scope: A two-stage framework that compresses a pretrained video autoencoder's latent — from f8t4p2 to f16t8p2 — while keeping it compatible with the pretrained DiT, reducing token count by nearly 8× and latency by 11.1× at 480×832×81 and 15.5× at 736×1280×81 on Wan2.1-14B.

What This Paper Is About

Video diffusion models are expensive because the Diffusion Transformer attends over a large number of latent tokens, and the token count is set by the autoencoder's compression ratio. Compressing harder naively hurts reconstruction, and compensating with wider latents both slows DiT convergence and shifts the latent away from the distribution the pretrained DiT already learned — forcing an expensive retrain or adaptation. GRACE's goal is to compress a pretrained video autoencoder while preserving the alignment between that autoencoder and its pretrained DiT, so the pipeline generates from far fewer tokens without losing generation quality.

Key Contributions

  1. A diffusion training framework (GRACE) that compresses a pretrained video autoencoder while preserving the pretrained autoencoder–DiT pipeline, explicitly narrowing the gap between reconstruction quality and generation quality.

  2. A dual-latent representation with asymmetric denoising. A base latent comes from the frozen pretrained encoder applied to a downsampled input, and a residual latent (C′=16 additional channels) carries the spatial detail and inter-frame motion the base cannot represent. During generation the base is denoised ahead of the residual by an offset δ.

  3. A generation-aware alignment objective (L_align) that matches the compressed latent to the pretrained latent inside the feature space of the frozen pretrained DiT, using cosine similarity over the first 10 of 40 DiT blocks, with a weight set adaptively from the ratio of reconstruction and alignment gradient norms.

  4. Demonstrated compression and speedups: nearly 8× fewer tokens, compressing both axes at once (16× spatial and 8× temporal), with an 11.1× latency reduction at 480×832×81 and 15.5× at 736×1280×81, while matching the pretrained pipeline's VBench quality on I2V and exceeding it on T2V.

Main Findings

  • Token and latency reduction. GRACE reduces Wan2.1-I2V-14B's token count by nearly 8× and end-to-end generation latency by 11.1× at 480×832×81, growing to 15.5× at 736×1280×81. In Table 2, GRACE runs in 75.8 s (T2V) and 77.7 s (I2V) versus 851.5 s and 863.2 s for Wan2.1-14B; in Table 3, 215.6 s and 218.8 s versus 3361.3 s and 3396.8 s, with 10.1k latent tokens versus 77.3k.

  • Generation quality preserved or improved. At 480×832×81, GRACE scores a VBench-I2V total of 87.90 against 87.92 for the pretrained Wan2.1-14B, and a VBench-T2V total of 85.81 against 83.93. At 736×1280×81, GRACE reaches 87.84 I2V against 87.88 and 84.74 T2V against 82.96.

  • The T2V gain is semantic. GRACE leads the next best model by 5.28 on the VBench-T2V semantic score, and the ablation traces this to both stages: the dual latent alone keeps the semantic score at the pretrained level, while generation-aware alignment and the base-ahead offset raise it by 3.90 and 2.43.

  • Reconstruction fidelity does not predict generation quality. Step-Video-VAE reconstructs 1.91 dB above LTX-VAE but scores 3.01 lower on VBench-I2V, even at 4× the tokens. The single-latent baseline reconstructs 1.13 dB above GRACE-VAE (33.76 dB PSNR versus 32.63 dB) yet scores 1.46 lower on VBench-I2V (86.44 versus 87.90). GRACE-VAE reaches the highest generation quality among compressed autoencoders at the smallest token count, within 0.02 of the pretrained pipeline, while Video DC-AE and Step-Video-VAE use 2× and 4× more tokens and score 2.98 and 3.87 lower.

  • Alignment and dual latents improve the latent distribution. In the t-SNE uniformity analysis, the coefficient of variation, Gini coefficient, and normalized entropy all improve with the alignment loss; the PCA view shows the dual latent's components become less noisy, and the alignment loss makes them follow object regions and hold structure across frames.

  • Direct comparison to DC-Gen. DC-Gen adapts the same pretrained DiT at twice GRACE's token count. GRACE runs faster and scores higher on every VBench total, although DC-Gen is higher on quality scores at 736×1280×81 — the authors attribute this to DC-Gen producing more saturated samples that the VBench LAION aesthetic predictor scores higher. GRACE leads DC-Gen on the I2V score by 7.23 at 480×832×81.

  • Asymmetric denoising helps under large motion. Figure 4 shows the base-ahead schedule recovering detail on the subject's face that a shared schedule degrades where motion is large. At inference the offset is fixed to δ=0.15, and both parts are denoised in the same forward pass, so the number of function evaluations is unchanged.

  • Training-cost context. Adapting the DiT dominates training cost. A DiT trained from scratch for the same number of steps falls far behind on both VBench-T2V and VBench-I2V. Under a shared budget of 38.5 H200 GPU days, both the Video DC-AE latent and the single-latent baseline fall short of the pretrained pipeline; GRACE spends 8.5 days on the autoencoder and 30 on the DiT.

  • LTX-Video's limitation. LTX-Video scores 86.00 on the VBench-T2V human action dimension at both resolutions, against 96.00–100.00 for every other model (99.00 and 100.00 for GRACE), and its I2V motion tends to come from a global zoom or slow camera movement while the scene stays static.

Methodology in Plain English

The authors start from a released, well-matched pair: a pretrained video autoencoder and a pretrained Diffusion Transformer that was trained on that autoencoder's latent space. Their premise is that this pair is already aligned, and that compression should be done in a way that preserves that alignment rather than breaking it.

Stage 1 — compressing the autoencoder. Instead of adding compression blocks and retraining only for reconstruction, they build a latent with two parts. The base part comes from the original frozen encoder applied to a spatially downsampled and temporally subsampled video, so it stays in the representation the DiT already knows. A second encoder, initialized from the pretrained weights, takes the full-resolution video and produces a residual part with 16 extra channels that supplies the detail and motion the low-resolution base misses. The two are concatenated along channels. A decoder, fully fine-tuned with residual channels zero-initialized, reconstructs from both.

To stop the residual channels from drifting into a space the DiT cannot use, they add an alignment loss. They run both the pretrained latent and the compressed latent — noised with the same timestep and independent noise — through the frozen DiT, trilinearly upsample the compressed branch's features to the full-resolution token grid, project them with a zero-initialized per-layer residual projection, and maximize cosine similarity to the pretrained features at the first 10 of 40 transformer blocks. The pretrained features are held fixed, so only the compressed representation is optimized. Because the two losses operate at different scales, the alignment weight is set automatically as the ratio of their gradient norms with respect to the residual encoder's last convolutional layer.

Stage 2 — adapting the DiT. The autoencoder is frozen. The DiT's input and output projections are extended from C to C+C′ channels and fully fine-tuned, while LoRA is applied to the transformer blocks. The distinctive step is asymmetric denoising: the base latent is kept ahead of the residual in schedule time by a randomly sampled offset during training, so the base is always the less corrupted of the two and acts as a stable anchor for the residual. The two timesteps share the timestep-embedding function but enter through separate modulation projections, with the residual projection zero-initialized and the pretrained modulation applied to the base, so adaptation begins from the pretrained behavior. The training objective is the average flow-matching loss over the two parts.

Evaluation. Reconstruction is measured on Panda-70M with PSNR, SSIM, LPIPS, and rFVD; generation is measured on VBench for T2V and I2V at 480×832×81 with 50 sampling steps and one video per prompt, with latency measured on a single A100 GPU. To compare autoencoders fairly for generation, the authors adapt the pretrained Wan2.1-I2V-14B to every latent under the same 38.5 H200 GPU-day budget.

Why This Matters

Impact on research. The paper argues against a common assumption that better reconstruction implies better generation, and provides a concrete mechanism — supervising the compressed latent in a frozen generator's feature space — for keeping a compressed representation usable by a generator that was never trained on it. The dual-latent plus asymmetric-denoising design offers a template for adapting pretrained generators cheaply rather than retraining them, and the reported result that Open-Sora 2.0's videos did not fully converge even on 160 GPUs after adapting a pretrained DiT highlights how significant the convergence problem is.

Real-world applications:

  • Faster, cheaper video generation services where inference latency and GPU cost dominate operating expenses.
  • Image-to-video and text-to-video content creation at 480p and 736p where a user iterates over many candidate clips.
  • On-device or edge deployment, where a nearly 8× token reduction directly lowers memory and compute requirements.
  • Long-video and high-resolution generation, where the paper shows the speedup grows with resolution (11.1× at 480×832×81 rising to 15.5× at 736×1280×81).

Industry relevance. The method reuses released pretrained weights rather than requiring new autoencoders trained from scratch, and its full training budget is reported in H200 GPU days (8.5 for the autoencoder, 30 for the DiT). That makes it attractive to teams that already run models such as Wan2.1 and want to cut serving costs without rebuilding their pipeline. The work was done while the first three authors were interns at Kakao Corp., with authors affiliated with KAIST AI and Kakao Corp.

Future Directions

  • Extending beyond a single pretrained family. All experiments use Wan2.1 (14B) as the pretrained pipeline, with LTX-Video and Open-Sora 2.0 appearing only as comparison points that use their own generators. Whether the approach transfers to other autoencoder–DiT pairs is not reported.

  • Compression beyond f16t8. The paper notes that the dual-latent representation alone is insufficient at higher compression ratios, where the residual must carry more information missing from the base. How far the residual-channel strategy can be pushed before alignment no longer suffices is an open question.

  • Evaluating quality metrics more robustly. The authors observe that VBench's quality scores favor the more saturated outputs of DC-Gen because of the LAION aesthetic predictor, and they report a human evaluation against Wan2.1-14B and DC-Gen in Appendix D.2. Building generation metrics that do not reward such artifacts is a natural follow-up.

  • Tuning the asymmetric schedule. Training samples the base-ahead offset from a uniform range and inference fixes δ=0.15. Whether per-content or per-step schedules help, and how the offset interacts with the number of sampling steps, is not established here.

Target Audience

Researchers and engineers working on video diffusion efficiency, latent representation learning, or autoencoder–generator co-adaptation. It is most useful to readers already comfortable with diffusion transformers, flow matching, and VAE-style latent compression, and to practitioners who want to shrink the serving cost of an existing pretrained video model without training a new one from scratch. Readers looking for an introductory treatment of video generation will find the method sections demanding, though the introduction and the experimental tables convey the central argument clearly.

Authors’ abstract

Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.

Read the original paper