Skip to content
AI.info

Research

XGait: A Multi-Modality Wireless Sensing Dataset for Indoor Human Tracking and Identification

XGait: A Multi-Modality Wireless Sensing Dataset for Indoor Human Tracking and Identification Overview Research area: Human-computer interaction and ubiquitous/mobile computing, specifically device-fr

arXiv
2608.07064
Published
2026-08-07
Authors
Wei Xu, Zhu Wang, Yifan Guo, Changlong Cheng, Yin Zhang, Zhihui Ren, Bin Guo, Zhiwen yu

AI summary

XGait: A Multi-Modality Wireless Sensing Dataset for Indoor Human Tracking and Identification

Overview

Research area: Human-computer interaction and ubiquitous/mobile computing, specifically device-free wireless sensing for indoor human tracking and identity recognition.

Technical level: Intermediate. The paper is a dataset and benchmark paper; the collection setup and dataset statistics are accessible to newcomers, while the Doppler-spectrogram pipeline, PLCR-based tracking, and the network candidates presuppose familiarity with wireless signal processing and deep learning.

Scope (one sentence): XGait is a synchronized Wi-Fi CSI, active acoustic, and camera-based dataset of walking from 27 participants across three indoor scenarios, released together with a unified Doppler-spectrogram pipeline and benchmark tasks for human tracking and identity recognition.

Venue and identifiers: Published in IMWUT, Volume 10, Issue 3, article 179, DOI 10.1145/3832029; arXiv:2608.07064v1 [cs.HC], 07 Aug 2026. Received 1 July 2026. License CC BY 4.0.

What This Paper Is About

Wireless gait sensing is attractive because it is device-free and preserves privacy, but a single wireless modality gives only a partial and indirect view of human motion, making it fragile to room layout, walking trajectory, and multipath effects. Existing public datasets compound the problem: they are mostly Wi-Fi-only, limited to simple or strongly constrained trajectories, and usually built for either tracking or identification rather than both. XGait addresses this by releasing synchronized Wi-Fi and acoustic walking recordings with camera-derived ground-truth trajectories, plus a shared representation and benchmark so that cross-modality complementarity can be measured rather than asserted.

Key Contributions

  1. The XGait dataset. A multi-modality wireless sensing dataset that synchronously captures multi-node Wi-Fi CSI and active acoustic signals, with vision-based measurements as ground truth for walking trajectories. It spans three scenarios and contains more than 22K walking samples from 27 participants, and the authors state it is the first large-scale dataset for multi-modality wireless sensing supporting both tracking and identification.

  2. A unified Doppler spectrogram representation. Heterogeneous Wi-Fi and acoustic signals are mapped into a shared time–frequency space using a velocity abstraction, harmonizing two physically different carriers (electromagnetic waves versus mechanical waves) for cross-modal comparison.

  3. A standardized benchmark pipeline. Pre-processing, temporal alignment (PLCR-driven soft alignment using torso motion as a modality-invariant anchor), and feature construction are standardized, enabling reproducible evaluation and systematic cross-modal fusion. Two tracking/recognition feature paradigms are supported: spectrum-based (deep learning directly on spectrograms) and model-based (structured motion descriptors such as PPVP).

  4. Quantified complementarity between modalities. Experiments report a task-oriented synergy: Wi-Fi provides superior tracking stability, while acoustics offers stronger biometric discrimination.

Main Findings

  • Scale and coverage: XGait contains more than 22K walking samples from 27 participants, spanning three indoor scenarios (laboratory, home, meeting-room) with LoS and NLoS conditions and diverse path geometries.

  • Valid recording counts: The dataset provides 22,288 valid Wi-Fi CSI recordings and 22,118 acoustic recordings as stated in the text; Table 2's total row lists 22,284 Wi-Fi samples. Vision provides 15,257 usable samples, because videos are occasionally unavailable or unreliable (limited field of view and occlusions in NLoS regions) and are excluded during post-processing.

  • Per-scenario sample counts (Table 2): Laboratory (20 participants): 10,816 Wi-Fi (43.7GB), 10,823 acoustic (18.8GB), 10,730 vision (240GB), 10,730 ground-truth trajectories. Home (19 participants): 4,753 Wi-Fi (20.5GB), 4,750 acoustic (8.6GB), 1,848 vision (33.7GB), 1,848 ground truth. Meeting-room (20 participants): 5,005 Wi-Fi (21.7GB), 4,835 acoustic (7.8GB), 2,000 vision (38.4GB), 2,000 ground truth. Meeting-room across different days (6 participants): 1,710 Wi-Fi (6.3GB), 1,710 acoustic (3.6GB), 679 vision (12.1GB), 679 ground truth.

  • Storage footprint: Totals listed in Table 2 are 92.2GB for Wi-Fi, 38.8GB for acoustic, and 324.2GB for vision, with 15,257 ground-truth trajectory files.

  • Participant attributes: 27 volunteers (4 female, 23 male) aged 20–35 years, with heights from 160 cm to 184 cm and weights from 55 kg to 105 kg.

  • Trajectory diversity is deliberate. The laboratory scenario defines nine unique path labels organized into five path clusters; path clusters 1–4 each contain both linear and curved trajectories, with linear segments labeled #1–#4 and their curved counterparts #6–#9, while path cluster 5 is a complex turning pattern labeled #5. All trajectories are bidirectional, with direction encoded by the parity of the repetition index.

  • Laboratory collection volume: 20 participants each complete all nine paths with 60 trials per path, yielding 540 trials per participant and 10.8K trials in total.

  • Home collection: Each path is repeated approximately 50 times by 19 volunteers, 15 of whom also appear in the laboratory dataset; accurate trajectory annotations are provided only for clusters 1 and 3 due to occlusion.

  • Meeting-room collection: 20 participants, of whom 13 also participated in the laboratory setting and 17 in the home setting, each completing 50 trials per path cluster; ground-truth trajectories are provided for clusters 1 and 2.

  • Intra-person variability is explicitly included. Six participants from the meeting-room group were recorded on multiple days over one week with naturally varying clothing, with the recording date encoded in the filename (for example, user1-1-1-d1-r1.dat).

  • Complementary physical strengths. The paper states that Wi-Fi typically offers broader coverage and stronger robustness in non-line-of-sight conditions, making it suited to coarse motion tracking, whereas active acoustic signals provide finer velocity resolution in line-of-sight settings, helping capture subtle gait dynamics.

  • Task-oriented complementarity. The evaluation reports that Wi-Fi provides a robust foundation for global trajectory tracking, while the finer spectral granularity of active acoustics yields more stable and discriminative signatures for identity recognition.

  • Comparison with prior datasets (Table 1). XGait is listed with 27 subjects, more than 22,000 samples, weak behavioral constraints, three modalities (Wi-Fi, acoustic, camera), and both tracking and identification functionality. Prior entries listed are Widar (5 subjects, 54 samples, 2017), CrossSense (100 subjects, 600,000 samples, 2018), Widar2.0 (6 subjects, 24 samples, 2018), GaitID (10 subjects, 4,600 samples, 2020), RISE (15 subjects, 1,150 samples, 2021), CAUTION (20 subjects, 900 samples, 2022), and EIGait (8 subjects, 1,920 samples, 2025). CrossSense's access is described as restricted.

  • Not reported in the available content: Quantitative accuracy, error, or F1 results for the tracking and identification benchmarks are not present in the truncated text provided, so no performance figures can be cited here. Details of Table 3 beyond the model identifiers M1–M5 are also cut off.

Methodology in Plain English

Collection hardware. The team built a heterogeneous sensing system connected over a local area network, with a software trigger starting and stopping all recordings together for each trial. For Wi-Fi, four mini-PCs with Intel 5300 network interface cards each connect to three external antennas spaced 2.5 cm apart; one node transmits and three receive. Nodes are mounted at 0.8 m height, the transmitter broadcasts IEEE 802.11n packets on channel 64 (5.31 GHz) at a 1 kHz sampling rate, and receivers capture CSI with the Linux CSI Tool, with each packet characterized by measurements from 30 subcarriers. Layouts are regular in the laboratory and meeting-room and irregular in the residential environment. For acoustics, four commodity audio devices are used: a JBL PS3500 speaker as transmitter and three UGREEN microphones (16-bit, 48 kHz) as receivers at the same height. Rather than the common monostatic setup, the authors use a bistatic multi-node configuration, emitting a single-tone continuous wave at 19 kHz and recording reflections at 48 kHz. For ground truth, a Canon EOS 80D DSLR captures 1080p video at 30 fps from a height of 1 m, with AprilTag fiducial markers for spatial calibration; trajectories are reconstructed by back-projecting per-frame foot-ground contact points (found via human pose estimation) into 3-D world coordinates.

Protocol. Each participant completes a 10–15 minute familiarization session before data collection. Participants respond to verbal "start" and "stop" cues; receivers and camera begin together on "start" and stop on "stop", with slight buffers before and after the walking interval to avoid truncation. Participants start and end at specified markers but keep their preferred walking speed and path curvature.

Unified representation. Because the two modalities differ in carrier frequency, bandwidth, sampling rate, and propagation, raw fusion is impractical. Wi-Fi CSI is pre-processed with CSI ratio and a 2–100 Hz band-pass filter. Acoustic echoes are down-converted to baseband by quadrature demodulation, low-pass filtered, and resampled from 48 kHz to 1 kHz to match the CSI packet rate; the strong near-DC component within ±15 Hz from static reflections and residual carrier leakage is then suppressed. Both streams are converted into Doppler spectrograms with the time–frequency reassignment spectrum (TFRSP), which concentrates energy better than standard STFT spectrograms. This yields 1 ms temporal resolution and 1 Hz frequency resolution.

Alignment. Because commodity devices lack a common clock, the authors use PLCR-driven soft alignment. The dominant Doppler ridge, tracked over time with dynamic programming enforcing local continuity, is converted to PLCR via PLCR = −f_D(t)·λ, where λ is the signal wavelength. One link (typically Wi-Fi) serves as the reference timeline; other links are shifted by the lag maximizing normalized cross-correlation between PLCR envelopes, then truncated or zero-padded. The motion window is found by thresholding the smoothed reference envelope, and all spectrograms are cropped to the minimum common length.

Tracking. Each link supplies a torso-level PLCR magnitude, but magnitude alone cannot tell whether the subject is approaching or receding, so a link-wise polarity sign is inferred from the energy asymmetry of positive and negative Doppler bins within the torso-dominant band. Signed PLCRs from all links of both modalities are stacked into one observation vector; the bistatic path length derivative is linearized so that signed PLCR approximates the projection of velocity onto a link-geometry direction vector. Velocity is solved by weighted least squares with a diagonal weight matrix based on link reliability, and position is updated with a discrete-time motion model.

Identity recognition. Two feature paradigms are benchmarked. Spectrum-based features feed time–frequency representations into neural networks per modality or jointly; for a 1 s gait segment the Wi-Fi and acoustic spectrograms have dimensions 201 × T_s and 801 × T_s respectively, with each link as a channel and samples trimmed or zero-padded to fixed length. Model-based features derive explicit motion descriptors; the authors note AcousticID is highly sensitive to trajectory and geometry, and that solving GBVP's high-dimensional optimization becomes computationally prohibitive for acoustic sensing due to the wider frequency range, so they adopt PPVP as the primary descriptor, producing features of size 360 × 40 × T_s with 1° angular resolution and 40 velocity bins. Five learning network candidates (M1–M5) are listed, ranging from a plain CNN with late fusion, through CNN+LSTM and CNN+Transformer variants, to a ResNet+Transformer model and a cross-attention Transformer with cross-attention fusion.

Why This Matters

Impact on research. The paper reframes single-modality wireless gait sensing as a fundamentally partial observation problem and supplies the data and pipeline to study complementarity directly. By pairing Wi-Fi CSI with active acoustics and vision ground truth, and by covering tracking and identification in one release, it enables joint evaluation from motion estimation through to biometric recognition, alongside controlled study of intra-person variability across recorded days and clothing.

Real-world applications (as motivated by the paper's stated application areas):

  • Indoor human tracking in homes and offices using existing commodity wireless infrastructure, without wearables or cameras.
  • Device-free identity recognition, where walking gait acts as a physiological and habitual biometric for recognition.
  • Health monitoring, since walking patterns reflect physiological and habitual characteristics and can be tracked unobtrusively over time.
  • Robust sensing in environments where vision is impractical, since wireless sensing offers better privacy and robustness to lighting conditions and visual occlusion.

Industry relevance. The dataset uses commodity hardware (Intel 5300 NICs, consumer speakers and microphones, a DSLR) rather than specialized laboratory equipment, and it follows Widar-style file naming to integrate with existing analytic pipelines and community tools. The bookkeeping also matters operationally: Wi-Fi data alone occupy 92.2GB while vision occupies 324.2GB, a practical signal about where the storage and privacy costs of such systems fall. The finding that Wi-Fi and acoustics fail in different conditions gives system designers a concrete argument for heterogeneous multi-node deployments, especially where asymmetric visibility across modalities occurs.

Future Directions

  • Deep fusion of trajectories and spectrograms. The paper explicitly identifies effectively fusing dynamic trajectory data with high-dimensional spectrograms inside a deep learning framework as an open issue; the spectrum-based paradigm currently operates largely without the subject's position.
  • Scaling model-based descriptors. GBVP-style body-centric velocity modeling is described as computationally prohibitive for acoustic sensing due to the wider frequency range, leaving a gap between what the model expresses and what can be solved in practice.
  • Generalization across environments and layouts. With regular layouts in the laboratory and meeting-room but irregular layouts in the home, and with deliberate LoS/NLoS variation including an obstacle placed near the center of the meeting-room sensing area, the dataset invites systematic study of how deployment geometry and environment specificity affect cross-modal fusion.
  • Biometric stability over time and population diversity. The multi-day recordings with varying clothing for 6 participants support analysis of wireless gait signature stability, and the participant pool of 27 volunteers (4 female, 23 male, aged 20–35) leaves room to test how findings extend to a broader population.

Target Audience

Researchers and practitioners in ubiquitous computing, wireless sensing, and human-computer interaction who work on device-free tracking, gait-based identification, cross-modal representation learning, or the construction and benchmarking of sensing datasets. It is most useful to readers who need a multi-modality dataset with vision ground truth and a standardized pipeline, and to those evaluating whether combining Wi-Fi with acoustic sensing improves robustness under challenging trajectories and propagation conditions.

Authors’ abstract

Wireless sensing has emerged as a promising approach for tracking and identification using commodity Internet of Things devices. However, the features derived from a single wireless modality are often fragile to variations in environmental layouts and walking trajectories. Furthermore, most existing studies are based on datasets collected in specific scenarios with limited trajectory diversity and sensing modalities, preventing a robust evaluation of system generalization. \textcolor{blue}{To address this gap, we introduce \textbf{XGait}, a multi-modality wireless sensing dataset that synchronously captures human walking using Wi-Fi and acoustic transceivers across three indoor scenarios, with vision-based measurements serving as ground truth. Specifically, XGait contains more than 22K walking samples from 27 participants, covering diverse directions and trajectories to support both indoor tracking and identity recognition. To bridge the heterogeneity of wireless sensing modalities, we propose a unified Doppler spectrogram representation that maps Wi-Fi and acoustic signals into a shared time--frequency space, along with a standardized benchmark pipeline for pre-processing, temporal alignment, and feature construction, enabling reproducible evaluation and systematic cross-modal analysis. Extensive evaluations demonstrate that Wi-Fi and acoustic sensing exhibit complementary strengths, particularly under complex trajectories and challenging propagation conditions, thereby paving the way for novel research in the field of multi-modality wireless sensing.} The dataset and code are available at https://github.com/warrior-087/XGait.

Read the original paper