Skip to content
AI.info

Research

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation

Overview Research area: Computer vision, specifically generative video models for long-form, multi-shot narrative video synthesis. Technical level: Advanced. The work builds on latent video diffusion

Multi-Grid Post-Training for Long-Form Multi-Shot Video Generation
arXiv
2609.06373
Published
2026-09-06
Authors
Jiawei Mao, Haoqin Tu, Hardy Chen, Yuhan Wang, Keyang Xu, Jieru Mei, Hongliang Fei, Ruogu Fang, Wei Shao, Cihang Xie, Yuyin Zhou

AI summary

Overview

  • Research area: Computer vision, specifically generative video models for long-form, multi-shot narrative video synthesis.
  • Technical level: Advanced. The work builds on latent video diffusion transformers, flow matching, 3D VAEs, LoRA post-training, and grid-structured latent representations.
  • Scope: The paper introduces MovieGrid, a post-training paradigm that rearranges temporally ordered video chunks onto a spatial grid so that a single generation can produce long multi-shot videos with consistent characters and environments.

What This Paper Is About

Existing video generators are biased toward preserving continuous motion, so when a whole multi-shot narrative is packed along one temporal axis ("Temporal Packing"), the model tends to collapse the story into a single continuous shot instead of realizing all the requested shots. The authors propose decomposing a long video into short ordered chunks, placing those chunks on a spatial grid, and jointly generating them so each local temporal axis handles far fewer shot transitions. The goal is long-form multi-shot video with coherent motion inside each shot and consistent characters, environments, and narrative progression across shots.

Key Contributions

  1. A grid-based reformulation of long-form multi-shot generation. MovieGrid transfers a factor of N (the number of chunks) from the temporal dimension to the spatial grid, representing all N × T frames using only T temporal steps without discarding any frames. This reduces the average number of shot transitions per temporal axis from 25.99 to 1.76 (reported as a 14.8× reduction).
  2. The Multi-Grid Long Video (MGLV) dataset. Constructed from 1,000 long-form source videos through a four-stage pipeline (Source Video Collection, Hierarchical Video Segmentation, Grid Video Construction, Character-Aware Story Annotation), yielding 54,281 grid videos paired with character-aware story prompts. A separate 64-grid counterpart contains 16,027 paired 64-grid videos and story prompts.
  3. Four grid-aware post-training components. Noise-Free Random-Grid Training (randomly keeping a subset of chunks clean as visual context), Grid Embedding (grid identity, grid geometry, and intra-grid position), character-aware Story Prompts with <grid N> and <C> tokens linking recurring entities, and Grid Boundary Loss (λ_GB = 0.1) supervising latent boundaries between adjacent grids.
  4. Demonstrated scaling along two axes. Increasing grid count from 16 to 64 extends a single generation from 1,616 to 6,464 frames under a fixed token budget, and successive grid videos can be conditioned on previously generated chunks for further extension without additional training.

Main Findings

  • More shots at the same token budget: Under the same token budget, MovieGrid generates 6.05× more video shots than the plain Temporal Packing baseline in a 1,616-frame video.
  • Shot detection improvement: Using TransNetV2 at a threshold of 0.5, the Temporal Packing baseline yields an average of 1.35 detected shots (78 of 89 outputs contained only one shot), whereas MovieGrid yields 8.17. Ordered Story-Shot Recall rises from 36.61% to 83.07%.
  • Intra-shot consistency (Table 1): MovieGrid scores 0.8970 for subject and 0.9291 for background, versus HoloCine at 0.7814 and 0.8358. The abstract separately reports an aggregate intra-shot consistency of 0.9131 for MovieGrid versus 0.8086 for HoloCine.
  • Inter-shot consistency (Table 1): MovieGrid scores 0.6139 for subject and 0.5689 for background, versus StoryMem at 0.5543 and 0.5224. The abstract separately reports an aggregate inter-shot consistency of 0.5914 for MovieGrid versus 0.5384 for StoryMem.
  • Competitive on other metrics: MovieGrid obtains aesthetic quality 0.5087, dynamic degree 0.7420, and semantic alignment 0.1959, which the paper describes as remaining competitive rather than state-of-the-art (HoloCine has aesthetic quality 0.5151, and StoryMem has dynamic degree 0.7573 and semantic alignment 0.2060).
  • Grid layout alone is not sufficient: A controlled Wan2.2-5B VIC-style baseline extended to 16 grids under matched LoRA and training settings produced incomplete grid structure. MovieGrid improved intra-shot consistency from 0.3362 to 0.9131 and inter-shot consistency from 0.2189 to 0.5914 relative to that baseline.
  • Length–resolution trade-off: MovieGrid-64 uses an 8×8 layout on a fixed 2560×1536 canvas, unpacking to 6,464 frames at 320×192 resolution. The paper reports a 13.45% drop across intra/inter consistency (75.23% to 61.78%) and only a 2.97% drop in semantic alignment (to 16.62%).
  • Component ablations (Table 2): MovieGrid reaches 0.9131 intra-shot, 0.5914 inter-shot, 0.5087 aesthetic, 0.1959 semantic. Removing Noise-Free Random-Grid Training causes the largest degradation, an average drop of 22.14% across the four metrics (55.23% to 33.09%). Grid Embedding causes a 13.26% average drop (55.23% vs 41.97%), Grid Boundary Loss 9.23% (46.00%), and character tags a smaller 3.00% drop (to 52.23%).
  • Long-range consistency case study: Qualitative results show MovieGrid preserving character identity and appearance across shot changes, non-human subject identity across distant shots, and fine-grained background details despite intervening content, where baselines exhibit identity drift.

Methodology in Plain English

The researchers first gathered 1,000 long-form YouTube videos lasting between 3 minutes and 4 hours, spanning cinematic, realistic, anime, cartoon, stop-motion, and 3D CGI styles. Each video was resampled to 30 FPS and cut into 1,296-frame subvideos, and each subvideo was divided into 16 non-overlapping 81-frame chunks. Crucially, these cuts are fixed temporal intervals, not detected shot boundaries, which avoids the cost and errors of shot-boundary detection; a chunk may contain one or several physical shots. The 16 chunks of a subvideo are then laid out in chronological order on a 4×4 spatial grid, producing one 81-frame grid video that represents all 1,296 original frames. Qwen3-VL 8B annotates each grid video in two stages: it first extracts timestamped character records and then uses those records to caption the corresponding time intervals, with the interval captions concatenated into a character-aware Story Prompt prefixed by <grid N> and using <C> tokens for recurring entities.

The model itself builds on the Wan2.2-5B backbone frozen during post-training, with rank-32 LoRA adapters and Grid Embedding modules optimized jointly—63.1M trainable parameters in total. Training uses 81-frame grid videos at a fixed 2560×1536 canvas (640×384 per grid) for 10 epochs on 8 NVIDIA B200 GPUs with AdamW, global batch size 8, learning rate 2×10⁻⁵, 100 warmup steps, and cosine decay. The key training trick, Noise-Free Random-Grid Training, activates with probability p_vis = 0.3; when active, a number N_vis drawn uniformly from {1, …, 8} of chunks are left noise-free and used as clean visual context while the rest are noised, and the loss is computed only on the noised chunks. This same mechanism allows conditioning a new grid video on chunks from a previously generated one. The total objective combines a standard flow-matching loss with the Grid Boundary Loss weighted by λ_GB = 0.1. At inference, the generated grid is unpacked and the chunks are concatenated along the temporal axis in grid-index order.

Evaluation uses 89 stories composed by GPT-5.6-Sol with no narrative overlap with the training set, spanning 3D CGI (20), anime (18), stop-motion (12), realistic (19), and cinematic (20) categories. Metric tooling follows VBench conventions: DINO and CLIP for intra-shot subject and background consistency, the LAION aesthetic predictor for aesthetics, RAFT for dynamic degree, and ViCLIP for semantic alignment; inter-shot consistency uses Grounding DINO and SAM to localize and segment prompted characters and environments, with DINOv2 measuring similarity of the same masked regions across shots.

Why This Matters

Impact on research. The paper challenges the assumption that a long multi-shot narrative must be laid out on a single extended temporal axis. By showing that a spatial grid plus noise-free context conditioning reduces per-axis shot load, it offers a concrete alternative to autoregressive extension, keyframe interpolation, and holistic joint generation, and it targets an under-explored supervision granularity: detector-free fixed intervals with aggregated character-aware story prompts rather than one-to-one shot prompts.

Real-world applications.

  • Automated production of multi-shot narrative content such as short films, episodic series, and animated shorts from a single story prompt.
  • Storyboarding and previsualization tools for film and advertising, where a script can be realized as an ordered set of shots with consistent characters.
  • Continuation of existing footage, since selected chunks from a prior generation can condition a new grid video to extend a story while preserving character identity.
  • Style-consistent content generation across the visual styles studied (3D CGI, anime, stop-motion, realistic, cinematic) for games, education, and social media.

Industry relevance. The method is a post-training procedure on an existing backbone (Wan2.2-5B) using LoRA adapters and a 63.1M-parameter trainable set, rather than a from-scratch pretraining run, which lowers the compute barrier for adopting it. The fixed-token-budget length scaling—from 1,616 to 6,464 frames by moving from 16 to 64 grids—is directly relevant to serving costs, and the ability to extend across successive generations addresses the practical need for arbitrarily long outputs. The work is partially funded by an unrestricted gift from Google, and the authors are affiliated with UC Santa Cruz, Google, University of Florida, and Vanderbilt University.

Future Directions

  • Mitigating the duration–resolution trade-off. MovieGrid-64 quadruples unpacked length at a fixed token budget but drops to 320×192 per-frame resolution; the paper reports this compromise without resolving it, leaving resolution-preserving long-video scaling open.
  • Extending successive-generation continuation. Cross-grid conditioning is demonstrated qualitatively for character-consistent continuation, but the paper does not report quantitative metrics on error accumulation or drift over many chained generations.
  • Replacing fixed-interval segmentation. Because the pipeline deliberately avoids shot detection and a chunk may contain multiple physical shots, comparing detector-free segmentation against learned or detector-based segmentation remains an unexplored ablation.
  • Broadening evaluation beyond the curated benchmark. The reported results come from 89 out-of-distribution stories authored by GPT-5.6-Sol across 5 categories; the paper explicitly notes that shot-oriented benchmarks such as ST-Bench evaluate a complementary setting with different supervision granularity, so generalization to other benchmarks and narrative domains is untested.

Target Audience

Researchers and engineers working on video diffusion models, long-form or narrative video generation, and post-training methods such as LoRA adaptation. It is also relevant to practitioners building multi-shot content pipelines who need character and scene consistency across shots, and to dataset builders interested in grid-based video representations and automated character-aware annotation. Readers should be comfortable with diffusion transformers, flow matching, latent representations, and standard video generation evaluation metrics.

Authors’ abstract

Generating long-form multi-shot videos requires coherent within-shot motion and visually consistent narratives across shots. Existing video generators favor continuous motion and struggle to present complete shot sets when an entire narrative is packed along one temporal axis. We propose MovieGrid, a Multi-Grid Post-Training paradigm that decomposes a long video into shorter, temporally ordered chunks and arranges them on a spatial grid for joint modeling. This design reduces the number of shots handled by each temporal axis while enabling global information exchange across chunks. We construct the Multi-Grid Long Video (MGLV) dataset from 1,000 long-form videos using source video collection, hierarchical segmentation, grid video construction, and character-aware story annotation, producing 54K grid videos paired with story prompts. Our Noise-Free Random-Grid Training retains a random subset of chunks as clean visual context for denoising the remaining chunks. Grid Embedding encodes grid structure, character-aware Story Prompts link recurring entities, and Grid Boundary Loss stabilizes layouts. Under the same token budget, MovieGrid generates 6.05 times more shots than Temporal Packing in a 1,616-frame video. On a benchmark spanning five real-world categories, it achieves state-of-the-art intra-shot consistency (0.9131 versus 0.8086 for HoloCine) and inter-shot consistency (0.5914 versus 0.5384 for StoryMem). MovieGrid can further scale video length with minimal compromise through single or multiple generations.

Read the original paper