Skip to content
AI.info

Research

HelixWorld: A Real-time Interactive Audio-Visual World Model

Overview Research area: Computer vision and generative world models, specifically real-time interactive audio-visual simulation (joint video and spatial stereo audio generation under user control). Te

HelixWorld: A Real-time Interactive Audio-Visual World Model
arXiv
2609.38123
Published
2026-09-29
Authors
Lei Ke, Jiahao Pan, Zeyue Tian, Jiaming Wang, Haoyuan Huang, Kam Man Wu, Pengjun Fang, Hongyu Liu, Chenyang Qi, Lin Wang, Ruibin Yuan, Weijia Chen, Fangneng Zhan, Qifeng Chen, Wei Xue, Yike Guo

AI summary

Overview

  • Research area: Computer vision and generative world models, specifically real-time interactive audio-visual simulation (joint video and spatial stereo audio generation under user control).
  • Technical level: Advanced. The paper assumes familiarity with diffusion/flow-matching models, distillation, autoregressive streaming, camera pose parameterization, and audio-visual representation metrics.
  • Scope: The paper introduces HelixWorld, an interactive world model that generates video together with camera-aligned spatial stereo audio in real time, along with a curated 4.1k-hour spatial audio-visual corpus, a causal distillation recipe, and a new benchmark (HelixBench) with 1,015 human-verified cases.

What This Paper Is About

Interactive world models let users steer generated video with keyboard actions or camera trajectories, but almost all of them produce silent video only. HelixWorld addresses the missing acoustic dimension: it couples visual scenes with camera-grounded spatial stereo sound that co-evolves under user interaction, so that sound pans and rotates correctly as the viewpoint moves. The goal is a system that both follows action controls and runs causally in real time, which the authors report as 24 FPS on a single NVIDIA H800 GPU.

Key Contributions

  1. HelixWorld model: An interactive audio-visual world model that natively unifies continuous 6-DoF camera trajectories and discrete actions, producing synchronized, camera-responsive video and stereo audio rollouts in real time.
  2. Causal distillation recipe: A framework combining causal initialization, an online trajectory distillation loss, and long-horizon streaming fine-tuning, which the authors report eliminates compounding rollout drift and sustains 24 FPS streaming.
  3. Multimodal data pipeline: A four-stage curation workflow that removes pseudo-stereo and non-diegetic audio tracks while annotating metric camera extrinsics and decoupled three-track captions (visual scene, foreground acoustic events, background ambiance).
  4. HelixBench: A benchmark of 1,015 human-verified interactive cases that formalizes spatial-acoustic consistency alongside temporal synchrony and multimodal semantics for interactive audio-visual evaluation.

Main Findings

  • Real-time efficiency: HelixWorld reaches a steady-state RTF of 0.77 at 768×512 and 24 fps on a single NVIDIA H800 GPU, including video and audio decoding. For comparison, the LTX-2.3 base model has RTF 14.61 and the only other audio-capable world model, EchoWM, has RTF 1.98 in the WBench navigation comparison.
  • Visual quality and control on WBench navigation split: HelixWorld scores 79.9 average, 79.2 quality, 76.5 setting, 86.4 interaction, 86.4 consistency, and 70.9 physical. Silent baselines scoring higher on average include Alaya-EVOKE-Turbo (82.0), EchoWM (81.0), and Zing-0.5 (81.0); HelixWorld is reported to outperform most open-source silent baselines on visual quality and control responsiveness.
  • Audio-visual quality on HelixBench (causal student): KL 1.3934, FAD 2.3872, DeSync 0.5867, ImageBind 0.2987, CLAP 0.3016, Spatial 41.7583. The bidirectional teacher scores KL 1.5414, FAD 3.0775, DeSync 0.4988, ImageBind 0.2700, CLAP 0.2676, Spatial 33.2428.
  • Cascaded dubbing fails at spatial alignment: Re-dubbing HelixWorld video with video-to-audio models AudioX, ThinkSound, and PrismAudio yields Spatial scores of −5.7392, 14.8467, and −9.5072 respectively. Negative values indicate the audio does not capture emitter locations, whereas joint generation achieves higher spatial alignment.
  • Temporal sync is not the strongest point: LTX-2.3 base reports the lowest DeSync (0.4042) versus HelixWorld causal student (0.5867) and EchoWM (0.6745).
  • Progressive conditioning helps camera control: Progressive fine-tuning reduces ViPE-recovered rotation error to 0.1144 and translation error to 0.0823, versus 0.1237 and 0.0957 for full fine-tuning.
  • Three-track captions help cross-modal consistency: Training on three-track captions (V + A + AV) and inferring on the same format gives the highest ImageBind similarity of 0.2462, compared with 0.2372 for video-only training and inference.
  • Trajectory distillation plus long-horizon tuning is best: On WBench, adding the trajectory distillation loss alone gives 78.7 average, long-horizon tuning alone gives 78.4, and both together give 79.1, versus 77.9 for neither. Long-horizon tuning alone improves quality to 79.0.
  • Trajectory distillation increases feature diversity: DINOv3 diversity rises from 0.1089 to 0.1326 (+21.7%) and CLIP diversity from 0.0718 to 0.0815 (+13.5%) when the trajectory loss is added to self-forcing.
  • Extended rollouts: In 30-second rollouts, trajectory guidance mitigates visual drift, maintaining cloud structures and scene landmarks with fewer artifacts than standard DMD.
  • User study: In a blind pairwise study with audio, HelixWorld is preferred over all baselines, winning the majority of votes in every matchup. With muted video, it outperforms Alaya, HY-WorldPlay, and SANA, ties with LingBot v2, and trails only EchoWM. Exact vote percentages are shown graphically in Figure 7 and are not reported numerically in the text.
  • Data curation yield: From 4.1k raw hours, the pipeline passes 3.4k hours after heuristic filtering (81.9% pass rate), 3.2k after semantic filtering (95.5%), and 3.0k after camera annotation (93.5%), yielding a final trainable corpus of 3.0k hours and 2.1M synchronized clips (73.1% overall yield).

Methodology in Plain English

The work proceeds in two training stages. First, a bidirectional teacher is trained on the curated corpus with two kinds of control signals: continuous 6-DoF camera poses (injected into visual attention via PRoPE) and discrete user actions from an 81-class vocabulary arranged over a 9×9 grid of translation and rotation bins (injected into the diffusion timestep embedding via AdaLN-Zero). Both controls enter the video branch and reach the audio branch through cross-modal attention, which is how camera ego-motion becomes stereo panning. Conditioning is introduced progressively in three steps (freeze the backbone and train control modules only, then unfreeze the video branch, then train the full audio-visual backbone) to avoid destabilizing pretrained representations.

Second, this teacher is distilled into a few-step causal student. Video latents and co-temporal audio tokens are grouped into temporal blocks with bidirectional attention inside a block and causal attention across blocks via a sliding key-value cache. The student is first initialized by regressing single-step predictions onto the teacher's Probability Flow ODE endpoints, then trained with self-forcing. Because plain distribution matching (DMD) alone causes mode collapse and long-rollout drift, the authors add an online trajectory distillation loss that supervises the student's velocity against the teacher's implied flow computed from intermediate rollout states, with stop-gradients and detached history. Each update randomly chooses the trajectory loss or the DMD loss. Long-horizon streaming tuning then extends rollouts segment by segment with a sliding KV cache that keeps sink frames as anchors and detaches distant history.

The data side is handled by a four-stage, cost-ordered pipeline: heuristic filtering (shot splitting, stereo probes, OCR-based overlay removal, perceptual quality scoring), clip standardization, semantic filtering with a multimodal LLM that removes non-embodied content and non-diegetic post-production audio, and parallel camera/caption annotation (VGGT-Ω poses with Depth Anything V3 metric scale, falling back to Metric3D v2; engine telemetry for game footage). The model builds on open-source LTX-2.3.

Why This Matters

Research impact: The paper argues that spatial-acoustic consistency, meaning whether synthesized sound fields physically track dynamic viewpoint motion, is a foundational criterion separating interactive world models from passive generation, and it provides the first benchmark dedicated to evaluating it. It also shows that naive causal conversion of joint audio-visual diffusion models produces catastrophic multimodal drift, and offers a distillation recipe to fix it.

Real-world applications:

  • Immersive VR/AR and spatial computing: Sound that rotates correctly with head and camera motion is a core requirement for believable virtual environments.
  • Game development and interactive media: Models like this could generate synchronized audio alongside procedurally generated scenes without hand-authored sound design.
  • Robotics and embodied AI simulation: Contact sounds carry cues about surface materials, contact forces, and off-screen events that are useful for training perception and planning.
  • Content creation and previsualization: Rapid prototyping of first-person or cinematic sequences with matched stereo audio.

Industry relevance: The reported 0.77 RTF on a single H800 GPU places real-time interactive generation within reach of practical deployment, in contrast to the 14.61 RTF of the LTX-2.3 base. The contributing organizations are The Hong Kong University of Science and Technology and Noiz AI, with a public project page (https://helixworld.org/) and GitHub repository (https://github.com/NoizAI/HelixWorld).

Future Directions

  • Improving temporal synchronization: The causal student's DeSync (0.5867) is higher than the LTX-2.3 base (0.4042) even though it achieves the best KL (1.3934), suggesting room to tighten audio-visual onset alignment.
  • Extending beyond the 30-second rollouts evaluated: The paper demonstrates drift mitigation in extended 30-s rollouts but does not report results at longer horizons or on memory/compute scaling beyond the four-block bounded inference context.
  • Scaling and diversifying the training corpus: The final trainable corpus is 3.0k hours and 2.1M clips, with 1.8k hours of real-world footage requiring estimated rather than ground-truth poses, leaving open how much of the spatial fidelity depends on engine-logged game telemetry.
  • Broadening the modality and evaluation scope: HelixBench covers 1,015 human-verified cases across six scene domains and four subsets; extending it to additional interactive modes, longer interactions, or other sensory channels are natural next steps. The paper does not report plans for these.

Target Audience

Researchers and engineers working on world models, video and audio diffusion, and interactive generative systems; practitioners building XR, gaming, and simulation pipelines who need synchronized spatial audio; and benchmark or evaluation researchers interested in spatial-acoustic consistency metrics. Readers without a background in diffusion distillation, camera parameterization, or audio-visual representation learning will find the methodological sections demanding.

Authors’ abstract

World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset with true stereo acoustics and metric camera poses, upon which we pre-train a bidirectional teacher conditioned on 6-DoF camera trajectories and user actions. To enable low-latency causal interaction, we distill the teacher into a few-step streaming student via an online trajectory distillation loss, sustaining drift-free joint audio-visual rollouts at 24 FPS on a single GPU. Furthermore, we formalize spatial-acoustic consistency and introduce HelixBench to evaluate whether synthesized sound fields faithfully track dynamic viewpoint motion. Extensive experiments demonstrate that HelixWorld matches state-of-the-art silent world models in visual fidelity and responsiveness, while significantly surpassing existing baselines in camera-aligned spatial-acoustic immersion.

Read the original paper