Skip to content
AI.info

Research

Realizing Immersive Volumetric Video: A Multimodal Framework for 6-DoF VR Engagement

Overview Research area: computer vision, volumetric video, VR/AR, multimodal 4D reconstruction, and spatial audio. Technical level: Advanced. This paper defines Immersive Volumetric Videos (IVV) and p

arXiv
2604.09473
Published
2026-04-10
Authors
Zhengxian Yang, Shengqi Wang, Shi Pan, Hongshuai Li, Haoxiang Wang, Lin Li, Guanjun Li, Zhengqi Wen, Borong Lin, Jianhua Tao, Tao Yu

AI summary

Overview

Research area: computer vision, volumetric video, VR/AR, multimodal 4D reconstruction, and spatial audio. Technical level: Advanced. This paper defines Immersive Volumetric Videos (IVV) and presents the ImViD dataset plus a multimodal pipeline that turns real-world multi-view video and audio into 6-DoF VR experiences.

What This Paper Is About

Most immersive VR content is computer-generated or falls short on full 6-DoF interaction, audiovisual feedback, high resolution, high frame rate, and long duration. The paper asks whether such immersive media can be built directly from real-world captured video and audio. It introduces a new media format, a capture dataset, and a reconstruction framework for producing temporally stable audiovisual volumetric content.

Key Contributions

  1. Defines Immersive Volumetric Videos (IVV) as a volumetric media format for VR/AR with large 6-DoF interaction spaces, audiovisual feedback, and high-resolution, high-frame-rate dynamic content. Introduces ImViD, a multi-view, multimodal dataset with 360-degree foreground-background capture, 5K resolution, 60 FPS, audio, and 1-5 minute scenes.
  2. Proposes a robust dynamic light field reconstruction framework based on a Gaussian spatio-temporal representation. It combines flow-guided sparse initialization, joint camera temporal calibration, and multi-term spatio-temporal supervision to handle complex motion and real-world capture imperfections.
  3. Presents a sound field reconstruction method for multi-view audiovisual data, described as the first of its kind for this setting. The approach is training-free for novel-view acoustic synthesis and supports moving sound sources.
  4. Establishes a unified pipeline for IVV production that jointly reconstructs dynamic light and sound fields. Validates the pipeline through benchmarks and immersive VR experiments.

Main Findings

  • Dataset scale and diversity: ImViD contains 7 scenes (5 indoor, 2 outdoor), 16 takes, 39 cameras, and 360-degree coverage. It provides 5K video at 60 FPS, static images at 5568x4176, total duration of 38 minutes 46 seconds, 139,560 frames per camera, and 2069.3 GB of data.
  • Capture strategy works: The mobile hemispherical rig captures complex indoor and outdoor scenes with rich foreground-background interactions. The authors introduce a spatiotemporal capture density metric, reporting about 0.10 m³/s for moving captures, compared with zero for static arrays and no captured volume for handheld monocular trajectories.
  • Reconstruction quality improves: The proposed Gaussian-based dynamic light field method surpasses existing methods on challenging in-the-wild motion, producing high-fidelity and temporally coherent 4D reconstructions.
  • Temporal calibration matters: Millisecond-level hardware synchronization is insufficient for fast motion. Learnable per-camera temporal offsets reduce blur and ghosting artifacts caused by sub-frame misalignment.
  • Audiovisual integration works: The sound field reconstruction module complements visual realism and supports moving sound sources. VR experiments demonstrate coherent 6-DoF audiovisual immersion with large interaction spaces.

Methodology in Plain English

The researchers built a custom capture rig with 39 GoPro cameras mounted on a hemispherical surface and placed on a remotely controlled mobile platform at human eye height. The cameras face outward to mimic natural human viewing. The rig also records synchronized audio. They use two capture strategies: high-density static image capture for the background, and dynamic video capture either from a fixed point or while the cart moves.

For visual reconstruction, they represent the scene as spatio-temporal 3D Gaussians. Each Gaussian stores position, shape, color, opacity, velocity, and timing information. Static parts are initialized once with long temporal extent. Dynamic parts are initialized per frame using optical flow to separate moving regions from static ones. This reduces redundancy and prevents floating artifacts.

They then jointly optimize the scene and learn a small temporal offset for each camera to correct sub-frame timing errors. The training uses several supervision signals: color matching, depth from a monocular depth model aligned to structure-from-motion points, optical flow matching to constrain motion, and perceptual loss. For audio, they reconstruct a spatial sound field from the multi-view audiovisual recordings. The method estimates sound direction and distance relative to the listener and generates binaural audio, supporting a moving sound source and a moving user.

Why This Matters

Research impact: The paper provides a benchmark and construction methodology for immersive volumetric video. It connects multi-view capture, 4D Gaussian reconstruction, and spatial audio into one pipeline, giving future work a shared dataset and baseline for real-world audiovisual 6-DoF media.

Real-world applications:

  • Virtual tourism and cultural heritage: users can walk through captured opera performances, classrooms, laboratories, and outdoor scenes with spatial audio.
  • Immersive training and education: realistic 3D recordings of teaching, meetings, and discussions can be explored from any viewpoint.
  • Holographic telepresence and remote collaboration: captured people and environments can be rendered as volumetric audiovisual content for more natural remote presence.
  • Entertainment and live events: music performances, sports, and interactive playback can be experienced with 6-DoF movement and directional sound.

Industry relevance: The work targets VR/AR headsets, volumetric streaming, telecom and media platforms, light-field displays, and metaverse applications. The involvement of Migu and China Mobile-related funding suggests direct interest in commercial immersive media and spatial audio services.

Future Directions

  • Scale the dataset and pipeline to more scenes, longer durations, and more complex acoustic environments with multiple or overlapping sound sources.
  • Improve compression, storage, and real-time streaming for 5K 60 FPS volumetric audiovisual content.
  • Handle reverberation, occlusion, and dynamic user head tracking more explicitly in sound field reconstruction.
  • Reduce capture complexity and cost, and test robustness across outdoor lighting, fast motion, and less controlled environments.
  • Integrate generative models or neural compression to broaden content creation beyond captured scenes.

Target Audience

Researchers and engineers in computer vision, computer graphics, VR/AR, volumetric video, 4D reconstruction, and spatial audio. Dataset builders, media production technologists, and advanced students working on immersive media will benefit most.

Authors’ abstract

Fully immersive experiences that tightly integrate 6-DoF visual and auditory interaction are essential for virtual and augmented reality. While such experiences can be achieved through computer-generated content, constructing them directly from real-world captured videos remains largely unexplored. We introduce Immersive Volumetric Videos, a new volumetric media format designed to provide large 6-DoF interaction spaces, audiovisual feedback, and high-resolution, high-frame-rate dynamic content. To support IVV construction, we present ImViD, a multi-view, multi-modal dataset built upon a space-oriented capture philosophy. Our custom capture rig enables synchronized multi-view video-audio acquisition during motion, facilitating efficient capture of complex indoor and outdoor scenes with rich foreground--background interactions and challenging dynamics. The dataset provides 5K-resolution videos at 60 FPS with durations of 1-5 minutes, offering richer spatial, temporal, and multimodal coverage than existing benchmarks. Leveraging this dataset, we develop a dynamic light field reconstruction framework built upon a Gaussian-based spatio-temporal representation, incorporating flow-guided sparse initialization, joint camera temporal calibration, and multi-term spatio-temporal supervision for robust and accurate modeling of complex motion. We further propose, to our knowledge, the first method for sound field reconstruction from such multi-view audiovisual data. Together, these components form a unified pipeline for immersive volumetric video production. Extensive benchmarks and immersive VR experiments demonstrate that our pipeline generates high-quality, temporally stable audiovisual volumetric content with large 6-DoF interaction spaces. This work provides both a foundational definition and a practical construction methodology for immersive volumetric videos.

Read the original paper