Research
Video Generation Models: A Survey of Post-Training and Alignment
Video Generation Models: A Survey of Post-Training and Alignment Authors: Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexi
- arXiv
- 2610.00812
- Published
- 2026-09-30
- Authors
- Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli
AI summary
Video Generation Models: A Survey of Post-Training and AlignmentAuthors: Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli (arXiv:2610.00812v1 [cs.CV], 30 Sep 2026; published in Transactions on Machine Learning Research)
Overview
- Research area: Post-training and alignment methods for video generation models within computer vision and generative AI.
- Technical level: Advanced — the paper assumes familiarity with diffusion models, transformers, latent spaces, reinforcement learning from feedback, and knowledge distillation.
- Scope (one sentence): The survey organizes all methods that adapt a pretrained video generation model after large-scale pretraining — without retraining it from scratch — into a taxonomy built on how alignment signals are enforced, and reviews the associated datasets, benchmarks, and open challenges.
Note on the available content: the paper text provided covers the abstract, introduction, the full taxonomy and survey structure (Sections 1–9), the preliminaries on generation settings, base models and alignment dimensions, and Section 3 through its discussion of domain adaptation. Sections 3.4 onward and the detailed contents of Sections 4–9 are not included, so specific benchmark scores, dataset sizes, and per-method evaluation numbers are not reported in the material available. The taxonomy in Figure 3 is referenced but not reproduced as text in the provided content.
What This Paper Is About
Video generation has moved from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics, but models trained only on large-scale web video still fail to reliably follow human intent, maintain temporal coherence, or respect physical and safety constraints. The paper argues that video alignment faces difficulties that image and text alignment do not — error accumulation over time, coupling between motion and appearance, trade-offs among conflicting objectives, and scarce supervision for temporal properties — which motivates systematic post-training rather than retraining from scratch. The goal is to give the first comprehensive review of post-training and alignment in video generation, organized by how alignment signals are applied rather than by architecture or conditioning type.
Key Contributions
- Post-training methods for video generation. A comprehensive review of post-training and alignment methodologies, covering supervised fine-tuning, preference- and reward-based optimization, self-training and distillation, and inference-time alignment and control.
- Taxonomy of alignment techniques. A structured taxonomy that organizes post-training approaches by their optimization mechanisms and alignment roles, with the key distinction being implicit versus explicit alignment, and with attention to video-specific challenges.
- Datasets and benchmarks for video alignment. A systematic summary of commonly used datasets, benchmarks, and evaluation protocols, categorized by alignment objective and temporal characteristics.
- A cross-family and multi-stage perspective. Section 7 compares the post-training families, examines how backbone architecture shapes post-training interfaces, and describes the multi-stage composition patterns used in modern video generation systems.
Main Findings
- Implicit versus explicit alignment is the organizing axis. The survey distinguishes methods that shape behavior indirectly through supervised adaptation, model-generated or teacher-provided signals, or structured controllability mechanisms (implicit alignment) from methods that directly optimize behavior using evaluative signals such as preference feedback, reward functions, or verifiable criteria (explicit alignment).
- Alignment is a spectrum, not a binary. Post-training methods differ in how directly, strongly, and reliably they influence aligned behavior; the implicit/explicit split is orthogonal to the training or inference mechanism used and reflects how alignment is enforced, not when or where optimization occurs.
- Four families cover the landscape. (1) Supervised fine-tuning (implicit), (2) self-training and distillation (implicit), (3) preference- and reward-based methods (explicit), and (4) inference-time methods (hybrid — they can enforce explicit alignment through evaluative guidance or support implicit alignment through iterative refinement and structured control).
- Pretraining optimizes likelihood, not human utility. Pretrained models tend to reproduce the "average" web video, including motion blur, static scenes, and uncurated compositions; post-training is framed as a shift in objectives from broad distribution modeling to targeted behavioral refinement.
- Failure modes are systematic. Reported failures include identity drift in long videos, physically unrealistic object interactions, unstable motion, and incomplete alignment with complex text prompts, reflecting a mismatch between pretraining objectives and real-world behavioral requirements.
- Four alignment dimensions. The survey groups objectives as instruction following and fine-grained controllability; temporal consistency and identity preservation; motion quality and physical plausibility; and aesthetic fidelity and safety — noting that these dimensions can conflict, creating stability-versus-expressiveness trade-offs.
- Pretrained models can be too conservative. Because limited motion reduces the risk of visible errors during generation, pretrained models often favor static scenes or very small movements; alignment in the motion dimension must counteract this.
- DiTs dominate modern systems. Latent Diffusion Transformers combining a spatio-temporal latent encoder (typically a 3D VAE) with a Transformer backbone, flow matching (often Rectified Flow), spatio-temporal patchification, and positional encodings such as 3D RoPE are described as the dominant paradigm in 2024–2025 systems, including Wan, HunyuanVideo, and OpenSora.
- Autoregressive models offer a different alignment route. Autoregressive approaches such as VideoPoet and VideoMAR formulate generation as next-token prediction over discretized spatio-temporal tokens, which permits direct application of established LLM alignment algorithms (for example standard PPO or DPO on token logits) without the adaptations continuous diffusion requires — yet the survey notes that most recent post-training advances and the open-source landscape center on diffusion, so the survey focuses on continuous diffusion trajectories.
- Parameter-efficient tuning is a recurring strategy. Methods such as LoRA and ControlNet attach to attention or modulation blocks in the DiT backbone, and analyses cited for panoramic video generation suggest low-rank updates suffice for that specialization.
- Research activity has grown rapidly. The survey reports that post-training alignment research for video generation has expanded rapidly since 2022, with growing diversity in supervision paradigms and deployment strategies (Figure 2 covers 2022–February 2026).
Methodology in Plain English
This is a literature survey, not an experimental study. The authors read across the video generation literature and built a framework rather than running new models.
The central move is a definitional one. They define post-training broadly as any optimization, adaptation, or control procedure applied after large-scale pretraining that changes a video generation model's behavior without retraining from scratch, and they define alignment as the degree to which a model's behavior conforms to desired objectives at deployment — following human intent, keeping temporal and identity consistency, respecting physical and causal constraints, and avoiding unsafe outcomes.
They then sort existing methods by the role of the signal used to shape behavior, which produces four families: supervised fine-tuning, self-training and distillation, preference- and reward-based optimization, and inference-time methods. Within this, they split methods into implicit alignment (behavior shaped indirectly, without explicitly checking whether outputs satisfy objectives) and explicit alignment (behavior optimized directly against evaluative signals), and they emphasize that this is a spectrum of strength rather than a strict partition.
To keep the review concrete, the authors also lay out the technical substrate that post-training must interact with — the T2V, I2V, and V2V problem settings, the components of latent diffusion transformer pipelines, and autoregressive token-based generation — because each component (latent compression, positional encoding, conditioning interfaces, flow-matching objectives) determines where post-training can intervene. The survey then walks through each family in a dedicated section, followed by cross-family comparison, datasets and benchmarks, and open challenges.
Why This Matters
- Research impact: The survey supplies a common vocabulary — post-training, implicit alignment, explicit alignment, and the four-family taxonomy — for a fast-moving literature that has grown rapidly since 2022 and has previously been organized mainly by architecture or conditioning signal. It positions itself as complementary to existing surveys on video diffusion models generally, controllable video generation, and human-centric video generation, which it says organize methods by architecture or conditioning type rather than by how alignment is enforced.
- Real-world applications (as domains discussed in the paper):
- Specialized fields such as healthcare and industrial physics, where general-purpose models degrade under domain shift.
- Long-form storytelling and scene-level content, where the requirement is cross-shot consistency rather than short, loosely connected clips.
- Instruction-guided video editing, where content is modified while preserving temporal coherence and physical plausibility.
- Personalized and identity-preserving generation, where characters or objects must keep the same appearance across movement, viewpoint change, and occlusion.
- Safety-aware deployment, where models must reduce harmful, biased, or NSFW output and refuse unsafe requests.
- Industry relevance: The surveyed methods target adaptation of existing pretrained systems rather than new architectures, with emphasis on parameter-efficient tuning, inference efficiency, and deployment-time control — the levers that matter when video generation models are expensive to train and to run. The paper's author list spans academia and industry organizations (Arizona State University, Twitch, Stanford University, eBay, NewsBreak, Microsoft, Columbia University, University of Southern California, Carnegie Mellon University), and a companion GitHub repository accompanies the survey.
Future Directions
- Scalable reward design. Building reward and evaluative signals for explicit alignment that scale, given that reliable supervision for temporal properties is scarce and expensive.
- Long-horizon temporal consistency. Handling error accumulation over time, identity drift in long videos, and temporal drift or positional mismatch that can arise from how token positions are represented.
- Stability-expressiveness trade-offs. Resolving the conflict between conservative, low-risk motion and the dynamic, physically plausible motion that alignment is supposed to encourage, within a fundamentally multi-objective problem.
- Safety-aware generation. Reducing harmful, biased, or NSFW content and enabling appropriate refusal behavior, alongside the related question of how to avoid bias introduced by proxy metrics or learned evaluators.
- Composition of multi-stage pipelines. Understanding how the different post-training families combine in modern systems, and how backbone architecture (diffusion versus autoregressive) shapes the available post-training interface.
Target Audience
Researchers and advanced practitioners in generative video who already understand diffusion models, transformers, and preference optimization, and who want a structured map of how to adapt a pretrained video model for controllability, consistency, physical plausibility, or safety. It is most useful to readers planning to build or select a post-training pipeline for a video generation system, to those looking for a taxonomy to situate new work, and to readers seeking pointers to datasets, benchmarks, and evaluation protocols for video alignment. Beginners will find the preliminaries section accessible but will need background in diffusion and reinforcement learning to follow the method families in depth.
Authors’ abstract
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherence, or satisfy physical and safety constraints. Compared with image and text generation, alignment in video generation presents unique challenges, including error accumulation over time, motion-appearance coupling, multi-objective trade-offs, and limited supervision for temporal properties. These challenges motivate systematic post-training strategies that adapt pretrained models without retraining them from scratch. In this survey, we present the first comprehensive review of post-training and alignment in video generation models. We frame post-training as a unifying framework and distinguish between implicit alignment and explicit alignment based on how alignment signals are enforced. From this perspective, we organize existing approaches into four broad categories: supervised fine-tuning methods, self-training and distillation methods, preference- and reward-based methods, and inference-time methods. This taxonomy provides a coherent view of how alignment signals shape model behavior across both training and deployment. Beyond methodological advances, we review commonly used datasets, benchmarks, and evaluation practices, and discuss open challenges such as scalable reward design, long-horizon temporal consistency, stability-expressiveness trade-offs, and safety-aware generation. This survey aims to provide a structured conceptual foundation and practical guidance for advancing controllable and reliable video generation models.