Skip to content
AI.info

Research

Multi-Frequency Federated Learning for Human Activity Recognition Using Head-Worn Sensors

Overview Research area: Federated learning (FL) for human activity recognition (HAR) using head-worn wearable sensors (earbuds and smart glasses), with a focus on devices that sample sensor data at di

Multi-Frequency Federated Learning for Human Activity Recognition Using Head-Worn Sensors
arXiv
2512.03287
Published
2025-12-02
Authors
Dario Fenoglio, Mohan Li, Davide Casnici, Matias Laporte, Shkurta Gashi, Silvia Santini, Martin Gjoreski, Marc Langheinrich

AI summary

Overview

Research area: Federated learning (FL) for human activity recognition (HAR) using head-worn wearable sensors (earbuds and smart glasses), with a focus on devices that sample sensor data at different frequencies.

Technical level: Advanced. The paper assumes familiarity with deep learning architectures, federated averaging, residual networks, and cross-validation protocols.

Scope: The paper proposes and evaluates a Multi-Frequency STResNet model that lets clients with sensors running at different sampling rates (simulated 5Hz or 3Hz low-battery modes versus 40Hz full-battery mode) jointly train a single federated HAR model on two head-worn datasets.

What This Paper Is About

Federated learning lets devices train a shared model without uploading raw user data, but existing HAR work generally assumes all participating devices sample sensors at the same rate. In practice, a device in low-battery mode may sample at 5Hz while others sample at 40Hz, and separate per-frequency models would waste data and shrink the client pool. This paper adapts the Spectro-Temporal Residual Network (STResNet) into a federated, multi-frequency setup for head-worn devices, so that clients sampling at different rates can contribute to one joint activity-recognition model.

Key Contributions

  1. A novel multi-frequency federated learning method for HAR based on STResNet, adapted to work in both federated and multi-frequency settings.
  2. Per-sensor spectral-temporal encoders tailored to each sensor channel's sampling frequency, combined with a context vector that masks activations from sensors that are not present on a given client.
  3. Evaluation against centralized and frequency-specific baselines on two head-worn datasets (USI-HEAR from earbuds and OCOsense from smart glasses), showing the multi-frequency model outperforms frequency-specific models and can use all available clients and sensors.
  4. A public implementation of the proposed network released on GitHub for further research.

Main Findings

  • STResNet was the strongest centralized model. On the USI-HEAR dataset with all sensor streams combined, STResNet reached 69.22% ± 11.78% accuracy, higher than ConvNet1 (57.12% ± 13.50%), ConvNet2 (65.49% ± 12.11%), ConvNet3 (61.78% ± 11.43%) and ConvNet4 (40.55% ± 13.25%). STResNet has 14,005,415 trainable parameters, second only to ConvNet1 (78,680,443 parameters) in computational demand.
  • Gyroscope was the most informative original sensor stream; derivatives beat magnitudes. Among virtual streams, first-order derivatives (DER) were a better input than magnitude (MAG) for most models.
  • Federated learning matched centralized learning. With all participants, FL scored 69.43% ± 4.14% accuracy on USI-HEAR versus 70.0% ± 4.7% centralized, and 85.19% ± 1.99% on OCOsense versus 84.9% ± 2.7% centralized. F1-scores were 67.26% ± 4.02% (FL) versus 70.2% ± 4.8% (centralized) on USI-HEAR, and 87.72% ± 1.72% (FL) versus 87.8% ± 1.7% (centralized) on OCOsense.
  • FL stayed robust as client counts dropped. With varying numbers of training participants (2, 3, 4, 6, 8), federated and centralized F1-scores remained comparable on both datasets.
  • The multi-frequency model beat single-frequency models. On USI-HEAR it reached F1 63.26% ± 3.08% at 5Hz and 65.53% ± 3.61% at 40Hz, against 59.30% ± 3.33% for exclusive 5Hz clients and 62.57% ± 4.49% for exclusive 40Hz clients. On OCOsense it reached 85.38% ± 2.52% at 5Hz and 83.45% ± 1.99% at 40Hz, against 72.81% ± 14.17% (exclusive 5Hz) and 79.74% ± 2.66% (exclusive 40Hz).
  • Downsampling everything to 5Hz was slightly better than the multi-frequency approach. The Down-5Hz baseline scored 65.38% ± 2.39% on USI-HEAR and 86.17% ± 2.29% on OCOsense, versus the ideal all-40Hz scenario at 69.14% ± 2.96% and 85.77% ± 2.17% respectively. The authors interpret this as a possible frequency threshold in these particular datasets.
  • On OCOsense, 5Hz was enough. Downsampling to 5Hz produced a higher F1-score (86.17% ± 2.29%) than the ideal all-40Hz scenario (85.77% ± 2.17%), suggesting 5Hz is sufficient for accurate HAR in that dataset.
  • The 3Hz "critical battery" experiments preserved the multi-frequency advantage. The paper states the multi-frequency model registered F1 of 80.69% ± 3.64% and 80.49% ± 2.51% for 40Hz and 3Hz, compared with 78.90% ± 3.01% and 66.26% ± 18.15% for single-frequency counterparts; Table VIII itself lists 78.96% ± 2.22% for the OCOsense 40Hz single-frequency row and 67.05% ± 15.89% for the 3Hz row.
  • The model tolerates a variable number of input sensors. Because of the context vector masking design, the authors report high flexibility and robustness while accepting a variable number of input sensors.

Methodology in Plain English

The researchers first compared five deep models under centralized training on the USI-HEAR dataset across several input configurations (accelerometer, gyroscope, magnitude, first-order derivatives, and all combined) to pick the best backbone. STResNet won, so it became the model used for all later experiments.

STResNet works by splitting each sensor channel into two pathways: a temporal pathway that uses residual blocks of 1D convolutions to capture motion over time, and a spectral pathway that converts the signal into a spectrogram and applies 2D convolutions. The two representations are fused and passed through dense layers to produce activity probabilities.

For federated learning, the authors used Weighted Federated Averaging implemented with the Flower library. Each client trains one epoch locally, then sends model weights plus its training sample count to the server, which computes a weighted average. Training used sparse categorical cross-entropy, an Adam optimizer with learning rate 0.0001, and a batch size of 32. Evaluation used person-independent 5-fold cross-validation with non-overlapping test clients across folds; the multi-frequency model was evaluated with person-independent 10-fold Monte Carlo cross-validation.

The multi-frequency extension gives each sensor channel its own encoder tuned to that channel's sampling frequency, then uses a context vector of 1s and 0s to mark which sensors are actually present, zeroing out activations for absent sensors before the fully connected layers. To simulate heterogeneous battery states, the authors treated half the clients as sampling at 5Hz (or 3Hz in a "critical battery" setting) and half at 40Hz, and compared the joint model against models trained only on 5Hz clients, only on 40Hz clients, all clients downsampled to 5Hz, and an ideal case where every client runs at 40Hz.

Why This Matters

Impact on research: This is described as the first study to explore a multi-frequency FL method for head-worn HAR in an asynchronous setup. Unlike approaches such as FLAME that require time-aligned synchronized devices from the same user, this method needs only one of the available devices or sensors to participate, which broadens who can join training and how much data the global model can use.

Real-world applications:

  • Chronic disease management, where activity tracking supports tailored healthcare interventions.
  • Healthcare monitoring that can automatically detect abnormal patient behavior and improve care efficiency.
  • Elderly care and personal fitness tracking, using unobtrusive earbuds or smart glasses rather than cameras.
  • Occupational health and safety programs that rely on continuous, low-power activity monitoring.

Industry relevance: Sampling frequency directly affects data stream size, storage, and compute load. A model that tolerates 5Hz or 3Hz clients alongside 40Hz clients lets consumer wearables contribute while in battery-saving modes, potentially reducing communication costs and enabling collaboration across heterogeneous hardware.

Future Directions

  • Unlabeled data. The authors did not leverage unlabeled sensor streams; they plan to explore unsupervised and semi-supervised methods to improve model quality and enlarge training data.
  • Personalization. No personalization techniques such as local re-training were applied; future work could use clustering or privacy-friendly knowledge sharing to address the non-IID problem in heterogeneous data.
  • Frequency thresholds and dynamic activities. Because downsampling to 5Hz was competitive on these datasets, the authors want to determine whether a frequency threshold exists, and expect 5Hz would not suffice for more dynamic real-world activities.
  • Communication and system efficiency. The authors frame multi-frequency tolerance in FL as a path toward lower communication cost, lighter client-side memory use, faster device response, and broader heterogeneous systems collaboration.

Target Audience

Researchers and practitioners working on federated learning, wearable sensing, and human activity recognition, especially those interested in earables, smart glasses, and resource-constrained or heterogeneous edge devices. It is also relevant to engineers building privacy-preserving health and wellness monitoring systems, and to readers who want a concrete baseline comparison between centralized and federated training on head-worn sensor data.

Authors’ abstract

Human Activity Recognition (HAR) benefits various application domains, including health and elderly care. Traditional HAR involves constructing pipelines reliant on centralized user data, which can pose privacy concerns as they necessitate the uploading of user data to a centralized server. This work proposes multi-frequency Federated Learning (FL) to enable: (1) privacy-aware ML; (2) joint ML model learning across devices with varying sampling frequency. We focus on head-worn devices (e.g., earbuds and smart glasses), a relatively unexplored domain compared to traditional smartwatch- or smartphone-based HAR. Results have shown improvements on two datasets against frequency-specific approaches, indicating a promising future in the multi-frequency FL-HAR task. The proposed network's implementation is publicly available for further research and development.

Read the original paper