Skip to content
AI.info

Research

Motion Prior Distillation in Time Reversal Sampling for Generative Inbetweening

Motion Prior Distillation in Time Reversal Sampling for Generative Inbetweening Authors: Wooseok Jeon (Yonsei University), Seunghyun Shin (GIST), Dongmin Shin (Yonsei University), Hae-Gon Jeon (Yonsei

Motion Prior Distillation in Time Reversal Sampling for Generative Inbetweening
arXiv
2602.12679
Published
2026-02-13
Authors
Wooseok Jeon, Seunghyun Shin, Dongmin Shin, Hae-Gon Jeon

AI summary

Motion Prior Distillation in Time Reversal Sampling for Generative Inbetweening

Authors: Wooseok Jeon (Yonsei University), Seunghyun Shin (GIST), Dongmin Shin (Yonsei University), Hae-Gon Jeon (Yonsei University, corresponding author) arXiv: 2602.12679v2 [cs.CV], 19 Feb 2026

Overview

Research area: Computer vision, specifically generative video inbetweening built on image-to-video (I2V) diffusion models and inference-time (training-free) sampling strategies.

Technical level: Advanced. The paper assumes familiarity with diffusion denoising, the EDM framework, classifier-free guidance, and time reversal sampling.

Scope: The paper diagnoses why time reversal sampling produces temporally incoherent inbetweening results and proposes Motion Prior Distillation (MPD), a training-free single-path sampling technique that distills the forward-path motion residual into the backward path.

What This Paper Is About

Generative inbetweening asks a model to generate plausible intermediate frames between a given start frame and end frame. Existing training-free approaches run two denoising paths — one conditioned on the start frame and one on the end frame — and either fuse them in parallel or alternate them sequentially, but each path carries its own motion prior, so the two paths disagree and produce ghosting, reverse-play motion, and other artifacts. The paper's goal is to eliminate this "motion prior conflict" by replacing the two competing priors with a single coherent motion prior taken from the start frame.

Key Contributions

  1. A formal diagnosis of bidirectional path misalignment. The authors recast existing time reversal sampling as an optimization problem (Eq. 11–12) whose loss minimizes the discrepancy between a path and its temporally reversed counterpart, and they show that incompatible motion priors from the two conditioning frames make this loss drive the sample toward worse, not better, trajectories.

  2. Motion Prior Distillation (MPD), a training-free sampling method. MPD computes the forward noise residual, then cumulatively subtracts it from the backward noise to reconstruct a denoised estimate that encodes the flipped motion prior of the start frame, deliberately avoiding any denoising of the end-conditioned path.

  3. Application to both parallel and sequential time reversal frameworks. MPD is plugged into TRF (parallel) and ViBiD (sequential), building on Stable Video Diffusion (SVD-XT), demonstrating that the idea is framework-agnostic.

  4. Quantitative, qualitative, and human evaluations. The authors benchmark on DAVIS (100 video-keyframe pairs) and Pexels (45 pairs) against six baselines, and run a user study with 30 participants over 28 randomly sampled video groups.

Main Findings

  • Bidirectional mismatch is the root cause of artifacts. The paper attributes ghosting, oscillations, reverse-play motion, and intermittent disappearance in prior methods to the forward-generation bias of I2V models — the backward path, initialized at the end frame, tends to generate forward-looking sequences rather than faithfully reconstruct historical frames. Figure 1 illustrates paths that "disagree on the car's destination."

  • MPD improves perceptual metrics on both benchmarks. On DAVIS, Ours + TRF reaches LPIPS 0.2212, FID 34.910, FVD 612.17, VBench 0.7992, VBench++ 0.9330; Ours + ViBiD reaches LPIPS 0.2220, FID 37.241, FVD 527.05, VBench 0.7845, VBench++ 0.9474. On Pexels, Ours + TRF reaches LPIPS 0.1149, FID 34.470, FVD 460.99, VBench 0.8503, VBench++ 0.9862; Ours + ViBiD reaches LPIPS 0.1028, FID 34.775, FVD 412.66, VBench 0.8235, VBench++ 0.9605.

  • Comparisons against six baselines. For DAVIS FVD, ViBiD scores 559.49, FCVG 621.82, GI 654.91, TRF 674.31, DynamiCrafter 678.92, and FILM 1058.0. For DAVIS FID, FCVG scores 38.997, ViBiD 39.883, GI 48.427, DynamiCrafter 46.739, TRF 56.894, and FILM 55.160. On DAVIS FVD, Ours + TRF (612.17) is higher than ViBiD (559.49), while Ours + ViBiD (527.05) is the lowest reported value.

  • FILM leads VBench++ on DAVIS. FILM scores 0.9740 VBench++ on DAVIS, above Ours + TRF (0.9330) and Ours + ViBiD (0.9474). The authors attribute this to flow-based warping preserving local structure near each endpoint, and note it comes with blurry artifacts and weaker long-range temporal consistency, reflected in FILM's FVD of 1058.0.

  • Human preference favors MPD. In the user study, ranking scores range from 3.5 to -3.5. Ours + ViBiD achieves the best alignment score (0.2440) with the lowest artifact rate (8.93%) and lowest unrealistic-motion rate (9.88%). Ours + TRF achieves alignment 0.3060, artifact 20.36%, unrealistic motion 22.62%. By comparison, TRF scores -0.3119 alignment with 28.09% artifact and 25.24% unrealistic; ViBiD scores -0.0678 with 28.10% and 25.24%; FILM scores -0.4060 with 62.74% and 54.76%; GI scores 0.1179 with 22.26% and 13.57%; FCVG scores 0.0988 with 20.36% and 19.17%; DynamiCrafter scores 0.0190 with 34.64% and 37.14%.

  • Distillation should happen early. Increasing the distillation step ratio γ consistently degrades results: γ = 0.2 gives LPIPS 0.2220, FID 37.241, FVD 527.05; γ = 0.4 gives 0.2478 / 45.636 / 634.11; γ = 0.6 gives 0.2562 / 55.873 / 813.08; γ = 0.8 gives 0.2679 / 68.007 / 973.64; γ = 1.0 gives 0.2721 / 75.544 / 1086.6.

  • Optimal hyperparameters differ by framework. For Ours + TRF, LPIPS and FID are minimized at γ = 0.3 (0.2212 / 34.910) and k = 2 (0.2212 / 34.910), while FVD prefers weaker distillation (γ = 0.1 gives FVD 573.13). For Ours + ViBiD, the optimum is γ = 0.2 (0.2220 / 37.241 / 527.05), k = 3 (0.2220 / 37.241 / 527.05), and λ = 1.0.

  • Modest inference overhead for no training. On a 25 × 1024 × 576 resolution setting, Ours + TRF takes 143 s and Ours + ViBiD 141 s at 19.2 GB VRAM, versus ViBiD at 108 s, FCVG at 134 s (24.1 GB), GI at 663 s (23.4 GB), TRF at 429 s (13.6 GB), and DynamiCrafter at 26 s (11.2 GB, 16 × 512 × 320). GI, FCVG, and DynamiCrafter all require training or fine-tuning; the proposed method, TRF, and ViBiD do not.

Methodology in Plain English

The starting point is a standard I2V denoiser (here Stable Video Diffusion) that, at each denoising step, predicts a clean estimate of the video. Time reversal sampling runs two such paths: a forward path conditioned on the start frame and a backward path conditioned on the end frame, combined either by linear interpolation at every step (parallel) or by alternating path segments with a re-noising step in between (sequential).

MPD replaces this two-path setup with a single, forward-driven path. During early denoising steps, the method measures how the forward denoised estimate changes from one frame index to the next — this is the "motion residual." It converts that residual into a noise residual, then starts from the end frame's latent and cumulatively subtracts the forward residual from it. The result is a reconstructed estimate that behaves as if the start frame's motion prior were playing backwards, which is exactly the motion the backward path should follow. Crucially, the end-frame condition is never used to denoise a separate path; the end frame only anchors the initialization. That reconstructed estimate is blended with the original forward estimate using an interpolation scale λ, and the blended estimate is plugged into the Euler update. The authors note this satisfies their original alignment objective in a relaxed form.

Two practical details matter. First, MPD is applied only in the early portion of sampling, controlled by a distillation step ratio γ, because early steps shape global, low-frequency motion while later steps refine high-frequency detail. After that window, sampling reverts to the existing time reversal sampler to tighten endpoint consistency. Second, a small number of extra re-noising steps k are used while MPD is active, and CFG++ is adopted following ViBiDSampler to limit off-manifold drift. Reported settings are an Euler scheduler with 25 timesteps, SVD-XT, and a single NVIDIA RTX 4090 GPU; the paper lists interpolation scale λ values of 1.0 and 0.5, re-noising steps k of 2 and 3, and distillation step ratio γ of 0.3 and 0.2 for the two configurations, and the ablation tables show the headline numbers for Ours + TRF correspond to λ = 0.5, k = 2, γ = 0.3, and for Ours + ViBiD to λ = 1.0, k = 3, γ = 0.2.

Why This Matters

Impact on research. The paper reframes a practical artifact problem — two denoising paths that disagree — as an optimization problem with a diagnosable cause, and addresses it without training or fine-tuning. Because MPD plugs into both parallel and sequential time reversal samplers, it is a general modification to a widely used family of inference-time strategies rather than a new architecture, which lowers the cost of adopting or extending it.

Real-world applications:

  • Frame interpolation for video post-production, where a large temporal gap between two shots needs plausible intermediate footage rather than blurry optical-flow warps.
  • Slow-motion and frame-rate conversion, where coherent motion direction matters more than per-frame sharpness.
  • Video restoration and archival upsampling, where missing or damaged frames must be filled consistently between two surviving frames.
  • Generative content pipelines for advertising, animation, or games, where a creator specifies first and last keyframes and expects a natural connecting motion.

Industry relevance. The method requires no training and no fine-tuning, keeps VRAM usage at 19.2 GB on the tested 25 × 1024 × 576 setting, and adds only a small inference-time penalty (141–143 s versus 108 s for ViBiD). Training-free deployment matters for organizations that cannot afford the fine-tuning budgets reported for GI (663 s inference after training) and FCVG (134 s after training), and the 30-participant user study provides the kind of human-preference evidence that production teams weigh alongside automatic metrics.

Future Directions

  • Reconciling the two framework optima. The parallel and sequential variants prefer different settings and show different metric trade-offs (Ours + TRF favors frame-level fidelity at some cost to temporal smoothness, per the authors' analysis). A unified criterion for choosing γ, k, and λ across frameworks is left open.
  • Extending beyond the early denoising window. The paper shows that applying MPD throughout sampling (γ up to 1.0) degrades results, but does not establish an automatic way to select where the distillation window should end.
  • Investigating the acknowledged VBench++ gap. FILM achieves the highest VBench++ on DAVIS (0.9740), and the authors attribute this to endpoint-local structure preservation. Whether MPD can retain endpoint fidelity while keeping its temporal coherence advantage is not resolved.
  • Testing on harder motion regimes. The datasets are chosen to include large motions such as driving and dancing, but the paper does not report behavior on cases where the start and end frames imply genuinely ambiguous destinations, which is where motion prior conflict is most severe.

Target Audience

Researchers and graduate students working on video diffusion models, frame interpolation, and inference-time sampling; practitioners who build on pre-trained I2V models such as SVD and need bounded (start-and-end constrained) generation without retraining; and engineers evaluating whether time reversal sampling is production-ready for video post-production and content generation pipelines.

Authors’ abstract

Recent progress in image-to-video (I2V) diffusion models has significantly advanced the field of generative inbetweening, which aims to generate semantically plausible frames between two keyframes. In particular, inference-time sampling strategies, which leverage the generative priors of large-scale pre-trained I2V models without additional training, have become increasingly popular. However, existing inference-time sampling, either fusing forward and backward paths in parallel or alternating them sequentially, often suffers from temporal discontinuities and undesirable visual artifacts due to the misalignment between the two generated paths. This is because each path follows the motion prior induced by its own conditioning frame. In this work, we propose Motion Prior Distillation (MPD), a simple yet effective inference-time distillation technique that suppresses bidirectional mismatch by distilling the motion residual of the forward path into the backward path. Our method can deliberately avoid denoising the end-conditioned path which causes the ambiguity of the path, and yield more temporally coherent inbetweening results with the forward motion prior. We not only perform quantitative evaluations on standard benchmarks, but also conduct extensive user studies to demonstrate the effectiveness of our approach in practical scenarios.

Read the original paper