Skip to content
AI.info

Research

ALERT Open Dataset and Input-Size-Agnostic Vision Transformer for Driver Activity Recognition using IR-UWB

Overview Research area: Wireless sensing and computer vision — specifically, Impulse Radio Ultra-Wideband (IR-UWB) radar for Driver Activity Recognition (DAR), combined with Vision Transformer (ViT) a

ALERT Open Dataset and Input-Size-Agnostic Vision Transformer for Driver Activity Recognition using IR-UWB
arXiv
2512.12206
Published
2025-12-13
Authors
Jeongjun Park, Sunwook Hwang, Hyeonho Noh, Jin Mo Yang, Hyun Jong Yang, Saewoong Bahk

AI summary

Overview

Research area: Wireless sensing and computer vision — specifically, Impulse Radio Ultra-Wideband (IR-UWB) radar for Driver Activity Recognition (DAR), combined with Vision Transformer (ViT) architectures and public dataset benchmarking.

Technical level: Intermediate to Advanced. The paper assumes familiarity with transformer architectures, radar signal representations (fast-time/slow-time, Doppler), and standard deep learning benchmarking practice, though the core problem is stated in accessible terms.

Scope: This paper introduces the ALERT open dataset of seven distracted driving activities captured with IR-UWB radar in real driving conditions, and proposes ISA-ViT, an input-size-agnostic Vision Transformer framework with range/frequency domain fusion for classifying those activities.

What This Paper Is About

Distracted driving causes a large share of fatal vehicle accidents, and IR-UWB radar is an attractive sensing modality for detecting it because it resists interference, uses low power, and does not capture images or audio. Two obstacles have blocked progress: no large-scale UWB dataset of diverse distracted driving behaviors collected in real (rather than simulated) driving environments, and the difficulty of feeding UWB radar data — which has non-standard, highly asymmetric dimensions — into ViTs that expect fixed-size image inputs. The authors address both by releasing the ALERT dataset (10,220 samples, 7 activities, real driving) and designing ISA-ViT, which resizes UWB inputs in an information-preserving way instead of naively interpolating them.

Key Contributions

  1. The ALERT dataset. Presented as the first UWB dataset capturing comprehensive driving activities in real-driving environments, with 10,220 samples across 7 activities and 9 volunteers. Benchmarking is provided using 8 learning algorithms spanning CNN-based, RNN-based, and transformer-based methods, and the dataset is released publicly on GitHub.

  2. The ISA-ViT framework. A model that adapts pre-trained ViTs to UWB data of varying input sizes. It resizes UWB inputs while preserving radar-specific information such as Doppler shifts and phase data, adjusts patch dimensions and linear projection layers to match altered patch sizes, and reuses pre-trained positional embedding vectors (PEVs) rather than cutting or interpolating them.

  3. A domain fusion strategy. A multi-view-learning-based approach that combines range-domain and frequency-domain (Doppler) representations to exploit their complementary signal characteristics and improve classification accuracy.

  4. Reported performance. ISA-ViT achieves a classification accuracy of 76.28%, a distracted driving detection accuracy of 97.35%, and 22.68% higher accuracy than the existing ViT method in UWB-based DAR, alongside benchmarking results and trade-offs for choosing models.

Main Findings

  • Dataset scale and composition: The ALERT dataset contains 10,220 samples, each 5 seconds long, collected from nine volunteers with heights from 165 to 188 cm, weights between 53 and 100 kg, ages 26 to 35 (average 32), and driving experience of 1 to 10 years (average 6). The seven labels are Relaxation (Relax), steering wheel control (Drive), nodding (Nod), smoking (Smoke), drinking (Drink), panel control (Panel), and using smartphone (Phone).

  • Preserving the pre-trained PEV sequence works better than manipulating it. In preliminary studies (Table V), "resizing input" achieved 52.20% accuracy on ALERT and 56.10% on RaDA; "adjusting patch shape" achieved 43.55% on ALERT and 53.92% on RaDA; "manipulating PEVs" (following the AST-style approach of Gong et al.) achieved only 36.74% on ALERT and 49.25% on RaDA.

  • Why manipulation fails: For ALERT and RaDA, the PEVs had to be interpolated for width by more than twice and 4.5 times, respectively. Excessive interpolation risks losing fine-grained spatial information or causing positional embedding saturation.

  • Why extreme patch shapes fail: Adjusting patch shape forces unconventional rectangular patches — 4×36 on ALERT and 74×2 on RaDA — which the authors report fail to capture sufficient information effectively.

  • Overall performance: ISA-ViT reaches 76.28% classification accuracy and 97.35% distracted driving detection accuracy, which the authors report as 22.68% higher accuracy than the existing ViT method for UWB-based DAR.

  • The simulation-to-reality gap is real. Citing Morales-Alvarez et al., the paper notes that recognition accuracy of driver observation models dropped from 85.7% to 46.6% when transferred from simulation to real-world scenarios, motivating collection in real vehicles.

  • ALERT versus RaDA: Compared with the prior RaDA dataset (6 activities, 10 subjects, 10,406 one-second samples, range-Doppler format, simulated driving, sensor placed in front of the driver with visual obstruction), ALERT offers 7 activities, 9 subjects, 10,220 five-second samples, both range-time and frequency-time representations, real driving, and air-vent sensor mounting. ALERT also uses more lenient activity restrictions.

  • Sensor and signal setup: Data was captured with the Novelda Xethru X4M06 radar: Tx center frequency 8.748 GHz, Tx bandwidth (−10 dBm) 1.5 GHz, energy per pulse 0.3–1.7 pJ, 100 frames per second, range bins 0–9 m (178 bins), sampling frequency 23.328 GHz, and pulse repetition frequency 15.1875 MHz. Two directional patch antennas operate between 7.25–10.2 GHz with a 65° opening angle in both azimuth and elevation.

  • Collection conditions: The urban route was 12 km, mostly flat with one uphill and one downhill segment of about 10° each, on asphalt with partial tiled-concrete. The campus route was a 6 km loop with a 30 km/h speed limit, roughly 50% sloped and 50% flat, with average gradients of about 10.2° uphill and 9.2° downhill, and included cobblestone, dirt, and speed bumps.

  • Benchmarking models: Eight algorithms were benchmarked — CNN-based (GoogLeNet, ResNet, DenseNet, MobileNet), RNN-based (a CNN+LSTM design with ResNet feature extraction over 10 snippets and a bidirectional LSTM with hidden state size 1024), and transformer-based (ViT and DeiT).

  • Fusion choice: Early fusion was selected for transformer models over late fusion, because late fusion would require two independent ViT models and significantly increase computational overhead, which is impractical for a real-time DAR system. The authors note late fusion achieves slightly higher accuracy but was not chosen.

  • Not reported in the provided content: The per-algorithm benchmarking results, the observation-time, multipath, and frequency-range analyses described as appearing in Section VI, and the detailed results tables are not included in the truncated text supplied here.

Methodology in Plain English

The researchers mounted an off-the-shelf UWB radar on a car's air vent — a position that keeps the driver's view clear, sits at roughly chest/eye level, and matches where drivers already clip phone holders. They drove two routes (a 12 km urban route and a 6 km campus loop) under speed limits with an experiment sign on the vehicle, and recorded seven activity classes from nine volunteers.

The raw radar returns are shaped by multipath reflections, producing a two-dimensional signal: a fast-time axis (range bins, i.e., distance) and a slow-time axis (successive transmitted pulses). This gives "range data." Applying an FFT to the range data yields "frequency data," which expresses Doppler frequency shifts and therefore motion characteristics. Both representations are provided in the dataset, along with 178 range bins and 178 frequency bins by default and an adjustable observation time window, so users can crop or resize inputs as needed.

For recognition, they benchmarked standard CNN, RNN, and transformer pipelines, adding a domain-fusion stage (concatenating range and frequency features) to each. The central methodological contribution concerns what happens when UWB data does not fit ViT's fixed 224×224 input and its 14×14 token grid. Rather than resizing the data (which discards detail), reshaping patches into extreme rectangles, or cutting/interpolating the pre-trained positional embeddings, ISA-ViT keeps the pre-trained 14×14 positional embedding sequence intact and adapts the input geometry in an information-preserving way, adjusting patch dimensions and the linear projection layers so altered patch sizes still work with pre-trained ViT weights. The model then fuses range-domain and frequency-domain features.

Why This Matters

Impact on research. The paper targets a concrete gap: prior UWB DAR datasets were collected in simulated driving, and prior ViT-on-UWB studies treated resizing as an implementation detail rather than a domain-specific design problem. By releasing a real-driving dataset with both range-time and frequency-time representations, and by showing quantitatively that positional embedding handling materially changes accuracy (52.20% vs. 36.74% on ALERT), the work gives the community both a benchmark and a cautionary result about transfer learning to non-image sensing domains.

Real-world applications:

  • In-vehicle driver monitoring systems that detect smoking, drinking, phone use, panel interaction, and nodding without cameras or microphones.
  • Monitoring of "relaxation" behavior during autonomous driving, addressing the paper's stated concern about blind trust in autonomy when drivers take their hands off the wheel.
  • Fleet, insurance, or commercial vehicle safety systems that need privacy-preserving sensing, since UWB captures no visual or audio data.
  • Smartphone-integrated sensing, since the authors note modern smartphones already include UWB chipsets and phone holders clip to the same air-vent location.

Industry relevance. The radar used (Novelda Xethru X4M06) is off-the-shelf and already widely used in UWB human activity recognition research. The 22.68% accuracy improvement over the existing ViT method, plus the deliberate choice of early fusion for real-time feasibility, frame the work as an engineering trade-off discussion relevant to automotive and mobile system designers, not only to academic researchers.

Future Directions

  • Testing generalization across subjects and vehicles. The dataset includes nine volunteers with a stated age range of 26 to 35; extending to broader demographics and vehicle types is a natural next question, and the paper's own comparison to RaDA (10 subjects, simulated driving) leaves open how models transfer across collection setups.

  • Scaling beyond seven labels. The authors argue that monitoring only one distracted activity is insufficient, and the ALERT labels represent one selection of fine-grained behaviors; adding more activities and more complex driving conditions is an evident extension.

  • Closing the remaining accuracy gap. Classification accuracy of 76.28% sits well below the 97.35% distracted driving detection accuracy, suggesting room to improve fine-grained discrimination between visually similar actions such as Smoke versus Drink, which the domain fusion strategy was specifically introduced to address.

  • Optimizing the fusion and efficiency trade-off. The paper notes late fusion achieves slightly higher accuracy than the early fusion it adopted for computational reasons, leaving open whether more efficient fusion architectures could recover that accuracy within real-time constraints.

Target Audience

Researchers and engineers working on wireless sensing, radar-based human activity recognition, and driver monitoring systems, particularly those using IR-UWB hardware or adapting vision transformers to non-image signal domains. It is also directly useful to practitioners who need a public benchmark dataset for in-vehicle DAR, and to automotive or mobile industry teams evaluating privacy-preserving sensing for driver-state monitoring. Readers without a background in radar signal processing will find the dataset description and benchmarking comparison accessible, while the positional embedding and patch-size analysis requires comfort with transformer internals.

Authors’ abstract

Distracted driving contributes to fatal crashes worldwide. To address this, researchers are using driver activity recognition (DAR) with impulse radio ultra-wideband (IR-UWB) radar, which offers advantages such as interference resistance, low power consumption, and privacy preservation. However, two challenges limit its adoption: the lack of large-scale real-world UWB datasets covering diverse distracted driving behaviors, and the difficulty of adapting fixed-input Vision Transformers (ViTs) to UWB radar data with non-standard dimensions. This work addresses both challenges. We present the ALERT dataset, which contains 10,220 radar samples of seven distracted driving activities collected in real driving conditions. We also propose the input-size-agnostic Vision Transformer (ISA-ViT), a framework designed for radar-based DAR. The proposed method resizes UWB data to meet ViT input requirements while preserving radar-specific information such as Doppler shifts and phase characteristics. By adjusting patch configurations and leveraging pre-trained positional embedding vectors (PEVs), ISA-ViT overcomes the limitations of naive resizing approaches. In addition, a domain fusion strategy combines range- and frequency-domain features to further improve classification performance. Comprehensive experiments demonstrate that ISA-ViT achieves a 22.68% accuracy improvement over an existing ViT-based approach for UWB-based DAR. By publicly releasing the ALERT dataset and detailing our input-size-agnostic strategy, this work facilitates the development of more robust and scalable distracted driving detection systems for real-world deployment.

Read the original paper