Research
ActionSplice: In-Flight Action Editing for Interactive World Models
Overview Research area: Generative video world models, specifically action-conditioned interactive video diffusion and inference-time state editing. Technical level: Advanced. The paper assumes famili

- arXiv
- 2609.08230
- Published
- 2026-09-08
- Authors
- Pardis Taghavi, Tingyu Guo, Jonas Lossner, Gaurav Pandey, Reza Langari
AI summary
Overview
- Research area: Generative video world models, specifically action-conditioned interactive video diffusion and inference-time state editing.
- Technical level: Advanced. The paper assumes familiarity with diffusion sampling, chunk-autoregressive latent video generation, classifier-free-style action conditioning, and solver-state resumption (clean-prediction versus Euler state transport).
- Scope: One-sentence scope: the paper proposes ActionSplice, an inference-time framework that edits the in-progress denoising state of a frozen, chunk-autoregressive video world model so that a new action received mid-sampling takes effect in the active chunk without waiting, without replaying completed solver evaluations, or without retraining the backbone.
What This Paper Is About
Chunk-autoregressive video world models generate a chunk of latent frames conditioned on a single action and denoise it through K solver steps. If a new action arrives partway through sampling, the partially denoised state has already been shaped by the old action, so current systems must either make the user wait for the next chunk, apply the new action only to the remaining solver evaluations (leaving a stale state), or roll back and redo the first r evaluations. The paper's goal is to make these models interruptible: incorporate an action change into the chunk already being generated, at the same solver step, while keeping the world model and sampler frozen.
Key Contributions
- Problem formulation as Counterfactual State Transport (CST): in-flight action editing is framed as transporting the interrupted, backbone-native representation to the matched counterfactual state that rollback under the revised action would have reached at the same solver step.
- Two editing variants: CST_R (retargeting) applies the revised action to the entire active chunk with mask M_0 = 1; CST_T (temporal splicing) preserves a temporal prefix and updates only the suffix from boundary m, using mask M_m.
- A lightweight learned corrector with matched supervision: a six-block residual 3D encoder-decoder with FiLM conditioning on camera controls and solver step, trained on matched rollback pairs (ordinary new-action rollback for CST_R, prefix-clamped rollback for CST_T), with recurrent training and rollback used only for supervision and evaluation references, not at inference.
- A matched evaluation protocol across two independently implemented backbones (minWM–Wan Action2V and HY-WM1.5), covering action-receipt steps, directed action transitions, temporal boundaries, repeated interruptions, human responsiveness annotation, and the HY-WorldPlay benchmark.
Main Findings
- Rollback-relative fidelity, whole-chunk: relative to direct condition swapping, CST_R reduces rollback-relative LPIPS by 61.5% on minWM–Wan Action2V and 75.9% on HY-WM1.5, and improves PSNR by 3.33 dB and 5.31 dB respectively.
- Suffix fidelity, temporal splicing: CST_T reduces suffix LPIPS by 56.1% on minWM and 77.5% on HY-WM1.5, and improves suffix PSNR by 3.30 dB and 9.52 dB relative to condition swapping.
- Latency and speedup: CST_T provides 2.73 times (minWM) and 1.69 times (HY-WM1.5) pixel-ready speedups over waiting; CST_R provides 2.73 times and 1.57 times. On minWM, CST_R reaches a pixel-ready latency of 2498.3 ms, essentially matching condition swapping at 2498.1 ms while achieving far lower LPIPS (0.1214 vs 0.3155).
- Boundary continuity: CST_R and CST_T obtain the lowest non-oracle boundary error on both backbones (CST_R: 0.0175 minWM, 0.0196 HY-WM1.5; CST_T: 0.0511 minWM, 0.0215 HY-WM1.5).
- Human-evaluated responsiveness on HY-WM1.5: CST_R reaches 94.7% action following with 1.05 mean stale frames, and CST_T reaches 100.0% action following with 0.40 mean stale frames; the Full Rollback oracle is 100.0% with 0.00 stale frames. Waiting is 21.1% / 15.37 stale frames (CST_R groups) and 26.7% / 7.53 (CST_T groups).
- Global video quality under HY-WorldPlay: CST_R obtains 25.66 PSNR, 0.6902 SSIM, and 0.1337 LPIPS against the original rollout, and 20.38 PSNR, 0.6267 SSIM, 0.1587 LPIPS in self-comparison — the best reported value for each metric in both comparisons (best prior comparison value reported: Light Interaction at 24.81 PSNR, 0.6500 SSIM, 0.1788 LPIPS).
- Camera-action following: CST_R reports R_dist 0.037 and T_dist 0.049; CST_T reports 0.035 and 0.120. Published short-term references listed are WorldPlay (0.031 / 0.121), CameraCtrl (0.037 / 0.341), VMem (0.048 / 0.219), and Matrix-Game 2.0 (0.287 / 0.843); the paper notes the request protocols differ, so these rows give context rather than a direct ranking.
- Repeated interruptions: a CST_R corrector trained on three-interruption trajectories and applied to five-interruption sequences (six action segments, all received after step r = 2 of K = 4) remains closer to sequential Full Rollback than condition swapping on every metric.
- Ablations: the full model performs best on every reported metric for both variants. Removing the hard output mask worsens CST_T prefix LPIPS (0.0412 to 0.0604) and boundary error (0.0215 to 0.0240). Unmatched supervision degrades every metric severely (CST_R action error 0.6192 to 4.6889, RB-LPIPS 0.0831 to 0.7797; CST_T prefix LPIPS 0.0412 to 0.4209).
Methodology in Plain English
The paper treats the internal state of an in-progress diffusion sampling run as something that can be nudged rather than restarted. When an action change arrives after r of K solver evaluations, the authors run a small learned network — the corrector — that looks at the interrupted state, the committed history, the old and new actions, and any retained sampler information, and predicts a residual. The residual is added to the interrupted state, optionally restricted by a temporal mask, and the result is inserted back into a valid solver state at the same step. Sampling then continues for the remaining K − r evaluations with the world model and sampler completely untouched.
The key training trick is matched pairing: the authors generate two branches that share the prompt, initial state, committed history, solver schedule, receipt step, and stochastic inputs, and differ only in the action schedule. One branch runs under the old action; the other runs rollback under the new action and provides the target. For CST_T, the target branch is prefix-clamped — after each replayed evaluation, positions before boundary m are replaced with the source branch's values — which makes the target correction exactly zero over the prefix. Training is recurrent, so correctors see histories produced by their own earlier predictions, and the loss is a masked normalized MSE on the residual plus a boundary-change term for CST_T.
Each corrector is a six-block residual 3D encoder-decoder with FiLM conditioning, trained separately per backbone and per variant. Backbone-specific handling differs: for minWM–Wan Action2V the corrector edits the cached clean prediction (pred_x0) and the native scheduler reconstructs the solver state with stored transition noise; for HY-WM1.5 the corrector edits the Euler solver state directly. Experiments use a fixed prompt-disjoint split of 150 scenes (120 training prompts, 30 held-out), three interruption captures per prompt per corrector (360 training, 90 held-out captures), balanced across six directed transitions among forward, backward, and yaw-left, with the main evaluation at receipt step r = 2 and a four-step sampler.
Why This Matters
- Impact on research: the work reframes control latency in interactive world models as a state-alignment problem rather than a scheduling or caching problem, and shows that frozen backbones can be made interruptible with a small, separately trained module. It also introduces matched counterfactual evaluation as a protocol for measuring in-flight editing fidelity.
- Real-world applications:
- Interactive game and simulation engines where a player's input must change the current frame, not the next chunk.
- Teleoperation and robot/vehicle simulators where a planner or operator issues mid-chunk steering corrections (the paper's forward, backward, and yaw-left action set).
- Camera-controllable content creation and virtual production, where a director changes camera direction while a shot is rendering.
- Streaming generative environments where a user prompt or control arrives asynchronously and stale frames are costly.
- Industry relevance: the method needs no backbone retraining, works on two independently implemented systems (minWM–Wan Action2V, HY-WM1.5), and its quantitative comparison includes deployed-style acceleration methods (TeaCache, SVG, BSA, Light Interaction) under the HY-WorldPlay protocol. Code and trained corrector weights are stated to be open source.
Future Directions
- Receipt-step generalization: the main results use r = 2 of K = 4; the appendix reports sweeps for r ∈ {1, 2, 3}, and how fidelity degrades at earlier or later interruptions is a natural extension (the appendix table was truncated in the provided content).
- Action-set generalization: the qualitative within-chunk example uses a yaw-right target that is explicitly noted as outside the corrector's training action set and is shown only as a qualitative generalization example, so systematic generalization to unseen actions remains open.
- Longer-horizon and higher-frequency interruption: the repeated-update experiment covers five interruptions on HY-WM1.5; whether transport error accumulates further over longer rollouts or with more frequent edits is not resolved.
- Broader backbone coverage: the paper trains a separate corrector per backbone and variant because the intervention structures differ, so whether a shared or transferable corrector across world models is feasible is left open.
Target Audience
Researchers and engineers working on interactive generative video, diffusion-based world models, and real-time streaming generation, particularly those concerned with control latency and inference-time state manipulation. It will also interest practitioners of diffusion sampling algorithms, inference acceleration, and human-in-the-loop interactive simulation, as well as readers who need a matched evaluation protocol for mid-generation control changes rather than end-of-rollout video quality. Prerequisite background in diffusion sampling and latent video generation is assumed.
Authors’ abstract
Chunk-autoregressive video world models typically condition each generated chunk on one action. An action received during sampling must therefore wait for the next chunk, condition future solver evaluations on a state produced under the previous action, or trigger rollback that repeats completed evaluations. We introduce ActionSplice, an inference framework that formulates this problem as Counterfactual State Transport (CST). A lightweight corrector transports the interrupted backbone-native representation toward the matched state induced by the revised action at the same solver step. The world model and sampler remain frozen, and sampling resumes without replaying completed evaluations. The retargeting variant $\mathrm{CST}*{R}$ updates the entire active chunk, while the temporal-splicing variant $\mathrm{CST}*{T}$ preserves a temporal prefix and updates only the suffix. Across minWM-Wan Action2V and HY-WM1.5, $\mathrm{CST}*{R}$ reduces rollback-relative LPIPS by 61.5% and 75.9% relative to direct condition swapping. $\mathrm{CST}*{T}$ reduces suffix LPIPS by 56.1% and 77.5%, respectively, while providing $2.73\times$ and $1.69\times$ pixel-ready speedups over waiting. Under the HY-WorldPlay protocol, $\mathrm{CST}_{R}$ obtains a PSNR of 25.66 dB, an SSIM of 0.6902, and an LPIPS of 0.1337 against the original rollout.