Research
WorldSonus: Bringing Sound to Worlds
Overview Research area: Audio generation, specifically video-to-audio (V2A) synthesis and spatial (stereo) audio generation for interactive video world models. Published under cs.SD (arXiv:2610.08760v

- arXiv
- 2610.08760
- Published
- 2026-10-06
- Authors
- Pengjun Fang, Jingyi Fa, Kam Man Wu, Jiaming Wang, Haoyuan Huang, Yaguang Wu, Xiangjun Huang, Ziyang Ma, Weijia Chen, Hongyu Liu, Zeyue Tian, Qifeng Chen
AI summary
Overview
Research area: Audio generation, specifically video-to-audio (V2A) synthesis and spatial (stereo) audio generation for interactive video world models. Published under cs.SD (arXiv:2610.08760v1, 06 Oct 2026).
Technical level: Advanced. The paper assumes familiarity with autoregressive transformers, diffusion/rectified-flow models, KV caching, contrastive learning objectives, and audio representation metrics.
Scope: WorldSonus is a modular, causal, streaming video-to-audio framework that generates real-time, text-controllable, spatially aligned stereo audio for interactive generative world models.
What This Paper Is About
Video world models can now synthesize realistic, interactive visual environments, but those environments are largely silent. Adding sound to them is hard because the audio must keep up with an interactive video stream in real time, must respond to text instructions that change mid-session, and must be stereo and spatially aligned with the scene geometry and camera motion. WorldSonus is a modular video-to-audio module that takes an externally produced visual stream and generates such audio causally, using only past and current video frames.
Key Contributions
- Real-time, interactive stereo generation. WorldSonus is a modular V2A framework that unifies real-time causal generation, mid-stream prompt responsiveness, and camera-aligned stereo synthesis for interactive world models.
- Causal audiovisual synchronization. The authors propose two-timescale visual conditioning (chunk-level visual summaries feeding the autoregressive backbone, frame-aligned local tokens feeding the flow head) plus a training-only ShiftNCE distillation objective that maintains temporal synchronization under causal streaming without adding runtime latency.
- State-of-the-art performance. Evaluations show competitive or superior acoustic quality and stereo balance relative to offline bidirectional baselines, with counterfactual tests confirming robust interactive control.
- A data curation pipeline for spatial audio. The authors assemble 1,465 hours of audio (999 hours of paired stereo video-audio plus 466 hours of audio-only data), including panoramic FOA recordings decoded into camera-aligned stereo, filtered in two stages with signal processing plus a Qwen3-Omni verifier.
Main Findings
- Real-time operation. WorldSonus generates 100 ms audio chunks with an execution latency of 41.2 ms per chunk on a single NVIDIA H100 GPU, giving a real-time factor of RTF = 0.41. Output is 48 kHz stereo audio.
- Acoustic quality. On 5 s clear-stereo VGGSound clips, WorldSonus achieves a VGGish FAD of 1.73. On 30 s interactive videos it achieves a FAD of 2.03. It reaches FAD 1.79 on VGGSound 10 s, 2.68 on Interactive 5 s, and 2.62 on Interactive 10 s, competitive with or better than the bidirectional baselines AudioX, ThinkSound, and PrismAudio.
- Advantage over the streaming baseline. Compared with the streaming monophonic baseline V-AURA, WorldSonus achieves lower distributional divergence and stronger audiovisual alignment while operating at a much finer temporal resolution (100 ms vs. 640 ms chunks).
- Spatial alignment. WorldSonus achieves the highest BiasSkill (stereo balance agreement) across all evaluated splits: 4.42 on VGGSound 10 s, 6.03 on Interactive 10 s, and 15.90 on Interactive 30 s. Qualitative spectrograms and energy-balance curves show it reproduces left/right channel dominance (for example left-side beach surf and a right-side passing truck).
- Temporal alignment. Bidirectional baselines that optimize directly on Synchformer features achieve lower DeSync, but WorldSonus stays competitive without test-time guidance. On the Greatest Hits onset benchmark it reaches accuracy 0.803, F1 0.785, and AP 0.871.
- Long-horizon stability. Over 1,024 Interactive 30 s clips, the rollout tail of an uninterrupted stream closely tracks the direct cold-start baseline (FAD 2.51 vs. 2.63; DeSync 0.827 vs. 0.880), indicating stable, collapse-free streaming under the bounded 5 s Ring-KV cache with no periodic resets.
- Interactive prompt control. On 10 s clips with sequential instructions, WorldSonus achieves dual relative match rates of 25.68% on VGGSound and 23.44% on Interactive, outperforming all baseline architectures. In paired counterfactual tests (switch vs. hold under identical visuals and seeds), it yields positive net intervention gains of +0.0509 and +0.0525 with 95% bootstrap confidence intervals excluding zero.
- Human preference. In a blind randomized A/B study on 40 Interactive 10 s clips, with twenty assessors providing 400 pairwise evaluations, overall preference reached 58.8% against AudioX, 80.0% against ThinkSound, and 65.0% against PrismAudio (ground-truth audio remained preferred over all generated results).
- ShiftNCE is important. Removing ShiftNCE degrades DeSync from 0.831 to 1.067, while direct feature injection from a 640 ms causal Synchformer window yields worse timing and acoustic quality.
- Flow-head visual conditioning is essential. Removing the direct frame-level visual path to the flow head causes FAD to rise from 2.42 to 15.82.
- DINOv3 beats pooled SigLIP 2. DINOv3 patch features improve ImageBind alignment from 24.56 to 28.05, and removing temporal deltas worsens DeSync (0.958) and FAD (3.08).
- Chunk size trade-off. At 33 ms, FAD worsens to 3.79; at 200 ms, latency increases without consistent gains; 100 ms was chosen as the practical balance.
Methodology in Plain English
WorldSonus treats audio generation as a streaming process over short chunks. Every 100 ms it receives three video frames (at 30 FPS) and the currently active text prompt, and predicts the matching audio latent chunk, which is three latent frames aligned with those video frames. Audio is generated in the continuous latent space of a frozen causal stereo VAE (adapted from SoundReactor) that encodes 48 kHz stereo into latents advancing at 30 Hz, with a stateful decoder that avoids re-decoding past latents.
The architecture combines a decoder-only autoregressive transformer with a rectified-flow head. The AR backbone keeps a bounded Ring-KV cache over a context window of W = 50 chunks (5 s), a circular buffer that overwrites the oldest slots so memory and compute stay constant as the stream grows. It processes two chunk-aggregated tokens per step: an audio summary from the previous chunk and a visual summary from the current chunk. The resulting hidden state conditions a compact flow head that denoises the current chunk's three latent frames, attending bidirectionally within the chunk only and using finer frame-aligned visual tokens.
Vision is handled at two timescales. A frozen DINOv3 encoder maps frames into spatial patch grids, each concatenated with its temporal difference to capture motion, then a learned query aggregates each frame's grid into one token. A second learned summary query compresses the three frame tokens into a single chunk-level token for the AR backbone, while the three frame-aligned tokens bypass the backbone and go directly to the flow head.
For text control, a T5Gemma 2 encoder and compressor turn the instruction into a compact token set conditioning the backbone by cross-attention, cached across steps. When the instruction changes mid-stream, the cross-attention cache is replaced in place at the nearest chunk boundary, preserving the Ring-KV cache and visual/audio state. Training includes time-varying prompt schedules with mid-sequence switches and drops.
Training has two stages: teacher-forced rectified-flow pretraining over a mixture of stereo video-audio and audio-only data (with Explorative Modeling evaluating three candidate noise draws per step and optimizing the best match), followed by fine-tuning on a high-quality stereo subset. A frozen Synchformer acts as a training-only teacher: ShiftNCE contrasts the aligned teacher embedding against temporally shifted embeddings from the same video, filtering out false negatives from static visual segments. The final loss combines the flow loss and the sync loss with a balance weight. Pretraining ran for 140,000 steps plus 10,000 fine-tuning steps, optimized with AdamW (weight decay 0.05). Baselines were run through their official inference pipelines with identical captions and chain-of-thought reasoning disabled.
Why This Matters
Impact on research. The paper shows that a bounded-memory causal model can match or exceed offline bidirectional models on acoustic quality and stereo balance, which challenges the assumption that full-clip visual context is needed for high-quality V2A. The training-only ShiftNCE idea is also transferable: it injects synchronization supervision from an external expert without paying inference latency. WorldSonus is deliberately modular, requiring only realized frames from an upstream world model rather than access to its action space or internals.
Real-world applications:
- Interactive games and gameplay video, where sound must follow player-driven camera motion and respond to changing scene instructions.
- Virtual and augmented reality, where stereo balance must track head and camera orientation for plausible spatial audio.
- Immersive film, simulation, and training environments needing synced sound effects for generated or edited footage.
- Any pipeline that generates video first and needs an add-on audio layer without retraining the visual backbone.
Industry relevance. The RTF of 0.41 on a single NVIDIA H100 means audio generation can run ahead of real-time video playback, which is the operating condition for live interactive products. The modular design lowers integration cost for teams that already have a visual world model but no audio module.
Future Directions
- Better causal audio codecs. WorldSonus relies on a frozen causal stereo VAE adapted from SoundReactor, which reports lower reconstruction quality for its causal decoder than for its non-causal counterpart; the authors identify more expressive large-scale causal audio codecs as key future work.
- More authentic stereo data. High-quality spatial audio remains scarce relative to monophonic corpora, and many public stereo videos contain artificial or noisy channel separation that does not reflect visual object motion, making dataset scale-up a stated goal.
- Closing the temporal-alignment gap. Bidirectional baselines still achieve lower DeSync on Synchformer measures, so improving causal event timing without test-time guidance is an open problem.
- Reducing dependence on external sync teachers and further validating interactive control. The paper's counterfactual prompt-switch tests isolate text responsiveness from visual transitions; extending this validation across longer sessions and more diverse prompt schedules is a natural next step.
Target Audience
Researchers and engineers working on video-to-audio generation, spatial audio, and generative world models, particularly those interested in streaming or causal inference constraints. It is also relevant to practitioners building interactive media, games, VR, and simulation systems who need a drop-in audio module for an existing visual generator, and to readers interested in training-time distillation of synchronization signals and in stereo/ambisonic data curation. The density of architecture details, metrics, and ablations makes it best suited to readers with prior exposure to diffusion and autoregressive generative modeling.
Authors’ abstract
Recent advances in world models have enabled increasingly realistic visual synthesis. However, these generated environments remain largely silent. Bringing sound to world models poses three core challenges: real-time generation to keep pace with interactive video streams, interactive control to respond to mid-stream sound instructions, and spatially aligned stereo to reflect scene geometry and camera motion. To address these demands, we introduce WorldSonus, an interactive video-to-audio framework designed for real-time spatial sound synthesis in world models. For real-time generation, WorldSonus employs a streaming causal autoregressive diffusion architecture that synthesizes audio chunks at a low real-time factor (RTF) of 0.41. For interactive control, we incorporate an audio-centric captioning pipeline with chunk-indexed prompt scheduling, enabling dynamic manipulation of sound events during generation. For spatial alignment, we leverage high-quality stereo supervision curated from diverse stereo and ambisonic data. Extensive experiments demonstrate that while tailored for world models, WorldSonus generalizes effectively to open-domain video-to-audio benchmarks, matching or outperforming state-of-the-art bidirectional models in both acoustic quality and spatial alignment. Project page: https://noizai.github.io/WorldSonus/