Research
Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training
Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training Overview Research area: Efficient attention mechanisms for diffusion-transformer-based joint video-audio gener

- arXiv
- 2610.05416
- Published
- 2026-10-04
- Authors
- Shuyuan Tu, Qi Tian, Yinming Huang, Yue Wu, Xintong Han, Kaihang Pan, Weijie Kong, Jiangfeng Xiong, Jian-Wei Zhang, Zuxuan Wu, Yu-Gang Jiang
AI summary
Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model TrainingOverview
Research area: Efficient attention mechanisms for diffusion-transformer-based joint video-audio generation, specifically sparse attention designed for native high-resolution (2K) training rather than inference-time acceleration.
Technical level: Advanced. The paper assumes familiarity with diffusion transformers (DiTs), block-sparse attention, flow matching objectives, and Triton kernel constraints on tile sizes.
Scope: The paper proposes Prism, a dynamic sparse attention framework that organizes tokens into spatiotemporal macro-zones, assigns each zone a content-adaptive 3D block shape, and applies hybrid Top-k/Top-p block selection to train joint video-audio generation models natively at 2K resolution.
Affiliations: Fudan University, Tencent Hunyuan Foundation Model Team, and Zhejiang University. arXiv:2610.05416v1 [cs.CV], published 04 Oct 2026, licensed CC BY 4.0.
What This Paper Is About
Training joint video-audio generation models at native 2K resolution would let them learn finer visual detail and sharper motion, but full attention scales quadratically with sequence length: a 2K clip (12 seconds, FPS=24) produces over one million tokens after VAE compression. Raising resolution adds mostly repetitive background tokens rather than informative ones, so full attention dilutes the weight placed on informative content and disrupts pretrained priors. Existing sparse attention methods either operate training-free or rely on fixed block shapes that ignore the fact that audio-visual coupling is concentrated in sound-producing regions. Prism's goal is to make native 2K joint video-audio training practical without sacrificing—and in fact improving—generation quality.
Key Contributions
-
An analysis of attention behavior in native high-resolution joint video-audio training, showing that full attention over-models redundant tokens at 2K and that recent sparse methods miss audio-visual coupling.
-
Prism, described by the authors as the first framework to explore native high-resolution training for joint video-audio generation, built on block sparse attention but replacing fixed block shapes with dynamic, content-adaptive ones.
-
A dynamic block shape mechanism that splits the video token grid into spatiotemporal macro-zones and estimates a per-zone variation indicator from two complementary signals: video channel-wise variance (how rapidly visual content changes along each spatiotemporal axis) and audio-to-video cross-attention norms (how strongly audio influences each visual region). A Lagrange-derived closed-form solution maps these signals to a block shape chosen from a candidate set of 16 shapes.
-
A hybrid Top-k and Top-p block selection strategy that determines per-query sparsity, where Top-k guarantees minimum coverage and Top-p adaptively expands the attended set for queries whose attention is broadly dispersed.
-
Experiments showing Prism trains 2.5× faster than full attention while surpassing it in generation quality, including on a purpose-built 2K benchmark of 300 unseen 2K videos.
Main Findings
-
Training speedup with better quality: Prism achieves a 2.5× training speedup compared to full attention while exceeding it in generation quality. The paper states this in the abstract, the introduction, and the contributions list.
-
2K-Bench results (native 2K): Prism outperforms all compared open-source methods on 2K-Bench. Its scores are AQ 0.61, DD 0.52, TF 0.982, ID 0.94, PQ 7.69, CU 7.34, DeSync 0.63, Sync-D 6.74, Sync-C 7.27, cpCER 0.187, MUSIQ 62.25, MANIQA 0.438, MotionQ 0.89. For comparison, MOVA scores AQ 0.42, TF 0.903, PQ 6.80, CU 6.26, DeSync 1.05, cpCER 0.374, MUSIQ 53.21, MANIQA 0.363, MotionQ 0.53; LTX-2.3 scores AQ 0.48, TF 0.943, PQ 7.05, CU 6.83, DeSync 0.95, cpCER 0.382, MUSIQ 55.60, MANIQA 0.403, MotionQ 0.66; Ovi scores AQ 0.38, TF 0.891, PQ 6.68, CU 5.92, DeSync 1.12, cpCER 0.468, MUSIQ 51.43, MANIQA 0.326, MotionQ 0.42; MagiHuman scores AQ 0.40, TF 0.912, PQ 6.72, CU 6.14, DeSync 1.08, cpCER 0.420, MUSIQ 52.68, MANIQA 0.337, MotionQ 0.58.
-
Full attention underperforms at 2K: MOVA, Ovi, and MagiHuman all employ full attention on the same native 2K data and yet underperform Prism, which the authors attribute to attention weight dilution, where massive redundant tokens claim most of the attention weight.
-
Smaller margins at 720p: On VABench and MOVA-Bench at the default 720p resolution (results reported in the paper's Table 3 and Table 4, whose numeric values are not included in the provided excerpt), Prism still outperforms all competitors but by a smaller margin, because shorter sequences contain less redundancy.
-
Ablation over sparse attention methods (Table 2): Full Attn scores AQ 0.42, TF 0.903, PQ 6.80, CU 6.26, DeSync 1.05, cpCER 0.374, MANIQA 0.363, MotionQ 0.53 at 0% base sparsity. Block Sparse Attention (BSA) with fixed shapes performs worst (AQ 0.33, MotionQ 0.34, 90% sparsity), as fixed-shape blocks mix semantically dissimilar tokens. VSA reaches MotionQ 0.46, SSTA 0.54, VMoBA 0.59, and SpargeAttn2 0.61. Training-free methods SVG-2 (MotionQ 0.38, 71% sparsity) and Sol-Attn (MotionQ 0.47, 85% sparsity) perform below Full Attn.
-
Improvement over the strongest sparse baseline: Prism improves over SpargeAttn2 by 46% in MotionQ and 30% in DeSync.
-
Dynamic block shape ablation: Fixed 8³ and Random Shape perform worst. The three-tier allocation surpasses the single finest tier (Only C₆₄) in both quality and speed. Removing the video channel-wise variance term (r_v,d) causes the most severe visual drop, while removing audio guidance degrades cpCER and DeSync, showing the two signals are complementary. Key features outperform Query features for variance estimation because Key encodes content closer to Value.
-
Shape mapping sensitivity: Performance peaks at the Lagrange-derived α = 1/2 and degrades on both sides: α = 1/3 forces blocks to remain nearly isotropic, while α = 2 is unstable under noisy variance. Alternative mappings (Softmax, Rank, Entropy) deviate from the optimal concave relationship, and an isotropic configuration confirms that anisotropic shape assignment is essential.
-
Threshold choice: τ₁₂₈/τ₂₅₆ peaks at (0.5, 0.25); tighter thresholds over-allocate fine blocks to backgrounds and looser ones assign coarse blocks to informative zones.
-
Resolution scaling: Full Attn leads at 480P where redundancy is minimal but drops at 2K, as redundant tokens overwhelm the learning signal. Prism is the only method that improves on all metrics from 480P to 2K. Removing video guidance causes visual metrics to basically stagnate beyond 1080P, while removing audio guidance fails to improve synchronization steadily. All competitors degrade on cpCER and DeSync at 2K.
-
Sparsity ablation (Table 7): Only Top-p lacks minimum coverage and loses critical context. Only Top-k at 85% achieves the best single-strategy quality, while 95% aggressively drops informative blocks and 75% admits more redundant tokens. The hybrid Top-k (95%) + Top-p (0.2) achieves the best quality-speed trade-off; Top-p (0.3) yields negligible gains but incurs 27% slower training time.
-
Backbone robustness (Table 6): The authors report that replacing video self-attention with Prism's dynamic sparse attention across different DiT backbones validates robustness, keeping all other components unchanged.
-
Comparison with commercial models: Commercial models outperform Prism in visual realism and audio aesthetics due to larger model parameters and datasets, but the gap is smaller in video-audio stability under large motions. Prism narrows the gap between open-source and industry-leading models.
Methodology in Plain English
Prism inherits the MOVA DiT and replaces only the video self-attention with sparse attention, since video self-attention dominates the cost at 2K while cross-attention and the audio branch operate on much shorter sequences.
Step 1 — Macro-zones. The video token grid of shape T × H × W is divided into non-overlapping macro-zones of shape Z_T × Z_H × Z_W, producing N_m = (T/Z_T)·(H/Z_H)·(W/Z_W) zones. Because the Triton kernel uses a fixed 64-token tile size and block sizes must be multiples of 64, the authors set Z_T/H/W = 8, which allows each zone to be tiled exactly by any 3D block shape whose per-axis edge lengths are in {2, 4, 8}.
Step 2 — Two guidance signals per zone. For each zone, Prism computes a variation indicator g_m = (g_T, g_H, g_W) independently per attention head and per layer.
- Video channel-wise variance guidance: The value features inside a zone are averaged along two axes at a time to produce axis-wise mean features, and per-channel variance along each axis is measured and normalized into an anisotropy ratio r_v,d. Value features are used because they are directly aggregated into the attention output. A large r_v,d means content varies rapidly along axis d, demanding finer partitioning there.
- Audio-to-video cross-attention norm guidance: Each video hidden state is modulated by cross-attention against audio keys and values. Prism measures the L2 norm of this audio injection per token, averages it per zone, and normalizes by the maximum zone average to map it into [0, 1]. It then computes per-axis variance of these norms to capture directional audio influence (for example, lip regions show high temporal audio variance as the mouth opens and closes). At the 2K scale this coupling is spatially compact.
The two are combined as g_d(m) = r_v,d(m) + ā(m)·v̂_a,d(m) + ε with ε = 10⁻⁴, where ā acts as a multiplicative gate: in sound-producing regions with high ā, block shape assignment is steered toward audio-visual coupling, and in silent regions the audio term vanishes.
Step 3 — Block size from information density. A raw information density ρ_raw(m) combines intra-zone spatial variance and temporal frame-difference variance, normalized to [0, 1] by dividing by the maximum over all zones. High ρ needs fine-grained blocks; low ρ is highly repetitive and benefits from larger blocks. With thresholds τ₁₂₈ and τ₂₅₆: ρ(m) ≥ τ₁₂₈ gives B(m) = 64, τ₂₅₆ ≤ ρ(m) < τ₁₂₈ gives B(m) = 128, and ρ(m) < τ₂₅₆ gives B(m) = 256.
Step 4 — Shape selection. For each size, the candidate set has 16 shapes total: C₆₄ contains 7 shapes, C₁₂₈ contains 6 shapes, and C₂₅₆ contains 3 shapes. Prism minimizes an intra-block information loss objective g_T·(b_T²−1) + g_H·(b_H²−1) + g_W·(b_W²−1) subject to b_T·b_H·b_W = B(m). Solving via Lagrange multipliers gives the closed-form b*_d ∝ g_d^(−1/2), so the axis with the largest variance receives the shortest edge. Because the solution is continuous while candidate edges are restricted to {2, 4, 8}, Prism quantizes by working in log₂ space, computing ideal log-edges ℓ*_d, and picking the candidate whose log₂ edge lengths are closest in squared distance. Token-level dynamic shapes within a zone are deliberately avoided because the Triton kernel requires memory-contiguous tiles.
Step 5 — Hybrid per-query sparsity. After partitioning, queries, keys and values share the same block layout and key blocks are mean-pooled into representatives. Block relevance is computed by dot products between each query block and key block representatives, followed by softmax, producing P_i over N_b key blocks. Top-k selects the k highest-weighted blocks, and Top-p selects the smallest set whose cumulative weight exceeds p. The union of the two sets forms the attended key/value blocks S_i, so that sharp queries select k blocks while flat queries may select more.
Training. Sparse attention is applied only to video self-attention. The model is trained with the video/audio flow matching loss, an expectation over noise levels of the squared error between the predicted velocity fields for video and audio and the true difference between clean and noise latents. Prism is initialized from MOVA and trained for 5 epochs at a learning rate of 1e-5, with Top-k (95%), Top-p (0.2), τ₁₂₈ = 0.5, and τ₂₅₆ = 0.25.
Data and evaluation. The training set is 100k 2K clips aggregated from UltraVideo and internet-collected videos. Evaluation is on MOVA-Bench and VABench, plus a new 2K-Bench of 300 unseen 2K videos (10 seconds long) with more complex motion patterns and appearance details, selected from the internet. VABench and MOVA-Bench experiments run at default 720p with Prism trained on the training set resized to 720p; 2K-Bench experiments train Prism and all competitors on native 2K data, and all models synthesize native 2K clips directly except LTX-2.3, which generates at 720p and upscales.
Metrics used: AQ, TF, DD, and ID for video quality; PQ, CU, cpCER, Sync-C, Sync-D, and DeSync for audio quality and video-audio synchronization; MUSIQ and MANIQA for high-resolution visual quality; MotionQ for overall motion quality.
Why This Matters
Impact on research. The paper reframes sparse attention for high-resolution generative training as a content-adaptivity problem rather than a pure efficiency problem. It argues that at 2K the bottleneck is not just compute but the dilution of attention weight by redundant tokens, and it introduces a design principle — make block granularity follow where content actually varies and where audio actually couples — that could transfer to other long-sequence multimodal diffusion settings. It also provides a negative result
Authors’ abstract
Natively training joint video-audio generation models at higher resolutions empowers them to learn richer visual details and sharper motion dynamics. However, full attention incurs quadratic cost and, as resolution increases, spreads attention over increasingly redundant tokens, diluting learning signals for informative content and disrupting pretrained priors. Existing sparse attention methods either target training-free acceleration or overlook the unique structure of joint video-audio data, where cross-modal interactions are inherently concentrated around sound-producing regions. To address this, we propose Prism, a dynamic sparse attention framework for natively training joint video-audio generation models at 2K. In particular, Prism organizes the token sequence into spatiotemporal macro-zones, enabling the attention structure to adapt to local content. For each zone, it estimates local information structure via video feature variance along the channel and feature norms from the audio-to-video cross-attention, jointly capturing how visual content varies directionally and how strongly audio influences each visual region. Based on these signals, Prism dynamically assigns a tailored block shape to each zone, applying finer partitioning along axes of rapid visual content variation and strong audio-visual coupling. This encourages tokens within each block to remain semantically coherent, allowing block-level features to capture both visual content and joint video-audio interaction patterns. Prism further adopts a hybrid block selection strategy to dynamically determine per-query sparsity. Experiments show that Prism achieves 2.5$\times$ training speedup compared to full attention, while surpassing it in generation quality.