Skip to content
AI.info

Research

Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy

Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy Overview Research area: Computer vision and audio-visual generative modeling — specifically methods

Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
arXiv
2609.34381
Published
2026-09-28
Authors
Abhinav Sharma, Sai Karthik Navuluru, Wang Wei, Daksh Dangi, Xiangbo Gao, Li Li, Bo Ni, Vardhan Dongre, Junda Wu, Xiyang Hu, Jiuxiang Gu, Seunghyun Yoon, Tong Yu, Chien Van Nguyen, Mohamed Elmoghany, Nedim Lipka, Hoda Eldardiry, Hongjie Chen, Tyler Derr, Thien Huu Nguyen, Zhengzhong Tu, Nesreen K. Ahmed, Franck Dernoncourt, Ryan A. Rossi

AI summary

Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy

Overview

Research area: Computer vision and audio-visual generative modeling — specifically methods that jointly generate, cross-generate, or jointly edit video and audio as a coupled pair.

Technical level: Intermediate. The paper is a survey and taxonomy. It is readable without deep mathematical background, but it assumes familiarity with diffusion models, latent autoencoders, and conditioning mechanisms.

Scope: The paper proposes a single formulation covering three problem settings (joint generation, cross-modal generation, joint editing) over one distribution on audio-visual pairs, and organizes the literature along a five-axis design taxonomy, with a first-of-its-kind taxonomy of joint audio-visual editing.

What This Paper Is About

Generative models for video and audio have largely been developed separately, so combining two strong unimodal systems naively produces streams that do not agree — a door closes but the impact sound is misaligned. The paper surveys the body of work that instead treats video and audio as a coupled pair, and asks one organizing question: how is the output kept coherent across modalities in time and semantics? It contributes a unifying formulation of the three settings, a five-axis comparison of methods, and the first systematic taxonomy of joint audio-visual editing.

Key Contributions

  1. First systematic taxonomy of joint audio-visual editing. The authors map editing of video and audio as a coupled pair — where an edit specified in one modality must propagate to the other — into nine categories and 28 edit types, each with representative operations and use cases (Table 4, Section 5).
  2. A unified formulation. Section 2 casts joint generation, cross-modal generation, and joint editing as three problems defined over a single distribution on audio-visual pairs, with stated definitions of audio-visual correspondence and latent representations.
  3. A five-axis design taxonomy. Table 1 organizes the literature along generation strategy, audio representation, video representation, alignment enforcement, and pretraining reuse, and marks how each covered method falls on each axis.
  4. Open problems grounded in the formulation. Section 13 states the open problems — long-horizon coherence, fine-grained control, physical plausibility, and evaluation — as instances of one underlying challenge: raising cross-modal alignment while preserving per-stream quality.

Main Findings

  • Three settings, one requirement. Joint generation outputs both streams from a condition, cross-modal generation outputs exactly one stream given the other, and joint editing outputs a modified pair given a clip and an edit instruction. All three must satisfy the correspondence requirement of Definition 1: agreement both semantically (sources visible are sources audible) and temporally (each acoustic event localized to the frames of its visual cause).
  • Editing adds two constraints beyond generation. Problem 3 requires propagation of an edit across modalities and preservation of content outside the targeted region. Editing therefore inherits the correspondence requirement and adds fidelity-to-input on top of it.
  • Dual-tower is the most common generation strategy. Two modality-specific towers exchange information through cross-modal connections; this is adopted by MM-Diffusion, JavisDiT, BridgeDiT, Ovi, UniAVGen, and LTX-2, among others. Single-tower designs are used by AV-DiT, MM-LDM, and recent open systems such as MOVA, Apollo, and 3MDiT. Cascaded generation is represented by Movie Gen.
  • Two taxonomy cells are unoccupied. No joint method covered by the survey adopts a unified-token strategy over discrete vocabularies, and none generates audio as codec tokens or directly in the waveform domain. UniForm comes closest to the unified-token cell — it serializes both modalities into one sequence and shares a denoiser — but operates over continuous VAE latents with a diffusion objective, placing it in the single-tower category.
  • Continuous latent audio dominates. Twenty-eight of the twenty-nine methods in Table 1 use a continuous latent audio representation with latent diffusion, spanning CoDi through LTX-2 and MOVA. MM-Diffusion is the representative mel-spectrogram instance.
  • 3D-VAE dominates video representation. Nineteen of the methods in Table 1 use a 3D autoencoder that compresses space and time jointly. CoDi, AV-DiT, SVG, and Animate-and-Sound instead use a per-frame 2D autoencoder plus a separate temporal module. MM-Diffusion is the sole pixel-space instance.
  • Cross-attention dominates alignment enforcement. It is used by nineteen methods, including MM-Diffusion, BridgeDiT, Ovi, and OmniForcing. Other mechanisms are shared positional encoding (ALIVE, JavisDiT++, AV-DiT, Apollo, 3MDiT), discriminator guidance (MMDisCo, the only instance), classifier guidance (Seeing-and-Hearing), and explicit prior (JavisDiT, retained in JavisDiT++).
  • Pretraining reuse splits three ways. Twelve methods train from scratch, including MM-Diffusion, Movie Gen, JavisDiT, Ovi, LTX-2, and MOVA. Single-pretrained reuse is used by ALIVE, Hallo-Live, and OmniForcing; dual-pretrained coupling by CoDi, Seeing-and-Hearing, SVG, BridgeDiT, and CCL. Guidance-based strategies necessarily entail dual pretraining, since they require two frozen unimodal models.
  • Wan 2.5 is partially closed. For this system, the alignment-enforcement and pretraining-reuse axes are left unmarked because the design is not publicly disclosed.
  • Datasets and metrics are covered but not specified in the available text. Sections 10 and 11 treat datasets, benchmarks, and metrics, but the truncated content does not report specific dataset sizes or metric values.

Methodology in Plain English

The authors did not run experiments. They reviewed the literature and built an organizing framework.

First, they wrote down a common mathematical setting: an audio-visual clip is a video (a stack of frames) paired with an audio signal over the same time interval, and methods learn a generative model over such pairs. On this basis they defined three problems — generating both streams, generating one from the other, and editing an existing pair — and argued that all three differ mainly in which signals are given and which are produced.

Second, they defined what "coherent" means, in two parts: semantic agreement (what you see is what you hear) and temporal agreement (each sound lines up with the frames of its visual cause). They note that a model producing good-looking video and good-sounding audio that score poorly on this alignment measure is perceived as broken even when each stream looks convincing alone.

Third, they proposed five independent design axes and placed each covered method along each axis using the published descriptions. The axes are complementary rather than orthogonal — a choice on one constrains choices on others. For example, a guidance-based strategy requires two frozen unimodal models and therefore entails dual pretraining, and any method reusing pretrained backbones inherits its representation from them.

Fourth, they reviewed the three settings in turn, described the building blocks methods inherit from unimodal work (VAEs, GANs, autoregressive models, diffusion and flow matching; video and audio representation spaces; conditioning mechanisms), and closed with cross-cutting sections on alignment and synchronization, architectures and training, datasets and benchmarks, metrics, applications, and open problems.

Why This Matters

Impact on research. The paper provides a shared vocabulary and comparison structure for a field that has been growing in isolated strands. By showing that joint generation, cross-modal generation, and joint editing are three problems on one distribution, it makes results comparable across settings. The marked empty cells of the taxonomy — no joint method on unified discrete tokens, no waveform-domain or codec-token audio generation — point at concrete unexplored territory.

Real-world applications (as framed by the paper's settings):

  • Soundtracking silent video. Cross-modal video-to-audio generation, where the footstep sounds of a person walking on gravel are synthesized and aligned to the frames where the foot contacts the ground.
  • Foley and sound effects. The survey treats Foley and sound effects as a distinct cross-modal task, with temporal onset as the dominant form of audio-visual agreement.
  • Dubbing and re-voicing. Joint editing with lip-sync as the primary correspondence, where changing the speaker's voice requires the lip motion in the video to be consistent with the new audio while background and identity stay intact.
  • Talking-head and speech-driven video. Generation where speech appears as both a condition and an output, with lip-sync as the dominating evaluation criterion.
  • Long-form and narrative editing. Listed among the joint-editing categories (Narrative) and joint-generation categories (Long-form), relevant to content production pipelines.

Industry relevance. The paper covers deployed and open systems from Adobe Research, Alibaba (Wan 2.5), Dolby Laboratories, Cisco, and others, and explicitly discusses pretraining reuse as a cost lever. A taxonomy that separates architecture choices from representation and alignment choices gives practitioners a way to reason about data and compute cost before committing to a design.

Future Directions

The paper states four open problems, all framed as instances of one underlying challenge: raising cross-modal alignment while preserving per-stream quality.

  1. Long-horizon coherence. Keeping audio and video in agreement over extended durations, where small offsets accumulate.
  2. Fine-grained control. Making conditions steer the output precisely across both modalities rather than coarsely.
  3. Physical plausibility. Making generated audio-visual events obey plausible physical relationships.
  4. Evaluation. Measuring cross-modal alignment and per-stream quality in a way that reflects what is actually perceived.

The taxonomy also raises structural questions: why unified discrete-token modeling is unoccupied among joint methods when token-based video and audio generation are established in the unimodal literature, and why no covered method generates audio in the waveform domain or as codec tokens.

Target Audience

Researchers and graduate students working on generative video, generative audio, or audio-visual learning who need a structured map of the coupled setting. Also useful to practitioners in media production, dubbing, and content generation who want to understand the design trade-offs — architecture, representation, alignment mechanism, and pretraining reuse — before choosing or building a system. The section on joint editing is likely the most distinctive entry point, since the authors state it is the first systematic taxonomy of that space.

Authors’ abstract

Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.

Read the original paper