Research
ActiTect: A Generalizable Machine Learning Pipeline for REM Sleep Behavior Disorder Screening through Standardized Actigraphy
ActiTect: A Generalizable Machine Learning Pipeline for REM Sleep Behavior Disorder Screening through Standardized Actigraphy Overview Research area: Clinical machine learning applied to wearable sens

- arXiv
- 2511.05221
- Published
- 2025-11-07
- Authors
- David Bertram, Anja Ophey, Sinah Röttgen, Konstantin Kufer, Gereon R. Fink, Elke Kalbe, Clint Hansen, Walter Maetzler, Maximilian Kapsecker, Lara M. Reimer, Stephan Jonas, Andreas T. Damgaard, Natasha B. Bertelsen, Casper Skjaerbaek, Per Borghammer, Karolien Groenewald, Pietro-Luca Ratti, Michele T. Hu, Noémie Moreau, Michael Sommerauer, Katarzyna Bozek
AI summary
ActiTect: A Generalizable Machine Learning Pipeline for REM Sleep Behavior Disorder Screening through Standardized ActigraphyOverview
Research area: Clinical machine learning applied to wearable sensor data — specifically, automated screening for isolated REM sleep behavior disorder (iRBD), a prodromal marker of Parkinson's disease and related α-synucleinopathies, using wrist-worn actigraphy.
Technical level: Intermediate. The paper combines signal-processing steps (resampling, filtering, auto-calibration, sleep–wake segmentation) with a standard gradient-boosted decision tree classifier, and evaluates it across multiple clinical cohorts. Readers should be comfortable with concepts such as cross-validation, AUROC, and feature engineering.
Scope: A single sentence: The paper introduces and validates ActiTect, an open-source, fully automated machine learning pipeline that turns raw multi-device wrist actigraphy into patient-level RBD risk scores, and demonstrates that it generalizes across four cohorts from three countries.
What This Paper Is About
Isolated RBD is an early warning sign of neurodegenerative disease that can precede motor or cognitive symptoms by up to 20 years, but the diagnostic gold standard — video-polysomnography — is expensive and requires expert manual scoring, while questionnaires are subjective. Wrist-worn actimeters could enable cheap, large-scale pre-screening by capturing abnormal nocturnal movements, but existing actigraphy analyses have been inconsistent, largely single-center, and often not released as usable tools. The goal of this work is to build a transparent, harmonized, end-to-end pipeline that works across different devices and acquisition settings, and to test honestly whether it holds up on independent, blinded, and externally collected data.
Key Contributions
- An open-source, fully automated end-to-end pipeline. ActiTect combines a non-ML standardization module (resampling, bandpass filtering, non-wear detection, auto-calibration, automated sleep–wake segmentation) with an XGBoost-based classifier, removing the need for manual sleep diaries or manual calibration.
- Device-agnostic harmonization with quantified quality checks. The pipeline's preprocessing steps are validated quantitatively: clock-drift correction, calibration error reduction measured against the unit sphere, filter passband retention/suppression, and sleep-detection agreement against both diaries and expert-scored PSG.
- Multi-center, multi-cohort validation including a blinded holdout and two external cohorts. Model development on one German trial cohort was tested on a blinded local set and on independent cohorts from Oxford (OPDC) and Aarhus (PACE), with leave-one-dataset-out cross-validation across all cohorts.
- A feature set designed with clinical domain knowledge. Features capture movement intensity, periodicity, spectral properties separating rapid from slow movements, complexity, fragmentation, overall activity, and clustering behavior — collaboratively designed with a sleep expert and checked for reproducibility across datasets.
Main Findings
-
Strong internal discrimination. On the 78-person development cohort (55 iRBD, 23 HC; 524 nights), nested cross-validation gave a patient-level AUROC of 0.95626, F1 of 0.92492 and balanced accuracy of 0.92333, versus night-level values of 0.87, 0.82 and 0.81 respectively — showing that aggregating across nights (6.4 ± 0.9 nights per patient) improves performance.
-
Well-calibrated probabilities. Brier scores were 0.14 at the night level and 0.12 at the patient level, indicating predicted probabilities reasonably reflect true RBD likelihood.
-
Blinded local generalization is moderate but clinically oriented. On the Local Test cohort (31 individuals, 198 nights), the final model reached AUROC 0.85533, F1 0.90115 and balanced accuracy 0.83267. The text notes this falls below the lower bound of the internal nested-CV 95% confidence interval (0.94–0.98), indicating a generalization gap, but recall (0.99) exceeded precision (0.83) — a favorable bias for screening, where false negatives are costlier.
-
External validation holds up across centers. On OPDC, iRBD vs HC gave AUROC 0.83772, F1 0.89026 and balanced accuracy 0.80999; PD+RBD vs HC gave AUROC 0.84363 but a lower F1 of 0.67124 due to class imbalance, with the combined all-RBD task at AUROC 0.83925. On PACE, iRBD vs HC was very strong (AUROC 0.97276, F1 0.95741, balanced accuracy 0.96065), while PD+RBD vs HC/PD–RBD retained high AUROC and F1 (0.95632 and 0.83591) but dropped in balanced accuracy (0.7765) — the paper attributes this to greater clinical heterogeneity and class imbalance shifting the optimal threshold.
-
Robustness across cohorts. Leave-one-dataset-out cross-validation across cohorts yielded an AUROC range of 0.84–0.89, and a complementary stability analysis reported that predictive features remained reproducible across datasets, supporting a pooled multi-center pre-trained model.
-
Auto-calibration is highly effective. Calibration error dropped from 50.89 ± 38.28 mg to 2.81 ± 0.81 mg (CogTrAiL-RBD), from 25.13 ± 7.65 mg to 3.09 ± 1.12 mg (Local Test), and from 40.01 ± 16.72 mg to 2.83 ± 1.01 mg (OPDC). Mean efficiencies (mean ± SD [95% CI]) were 0.93 ± 0.04 [0.92, 0.94], 0.87 ± 0.04 [0.85, 0.89], and 0.91 ± 0.05 [0.91, 0.92] respectively.
-
Automated sleep segmentation matches manual references. Against sleep diaries in CogTrAiL-RBD (61 individuals: 44 iRBD, 17 HC, 7 nights each), the HDCZA algorithm achieved a mean c-statistic of 0.93, mean absolute error of 34.4 minutes, and Pearson correlation of 0.994. Against expert-scored PSG in Local Test (16 participants: 13 iRBD, 3 HC), it achieved a c-statistic of 0.91, MAE of 35.8 minutes, and correlation of 0.996. The algorithm slightly underestimated sleep duration relative to PSG, which the authors suggest may be beneficial by reducing the chance of computing features on wake-period activity.
-
Filtering preserves signal and removes noise. A 0.8–20 Hz bandpass gave retention/suppression scores of 0.78 ± 0.01 [0.77, 0.78] / 0.89 ± 0.01 [0.89, 0.90] for CogTrAiL-RBD and 0.73 ± 0.09 [0.70, 0.77] / 0.90 ± 0.01 [0.89, 0.90] for Local Test.
-
Wrist dominance did not matter. Using 111 OPDC participants with simultaneous bilateral recordings and a paired permutation test (50,000 permutations), no statistically significant differences were found: AUROC dominant 0.85 vs non-dominant 0.88 (p = 0.28), balanced accuracy 0.82 vs 0.81 (p = 0.84), and F1 0.91 vs 0.90. The complete F1 p-value is cut off in the provided text.
-
A prior published approach transferred poorly. The authors first evaluated a previously published actigraphy-based RBD detection method (Raschellà et al., 2023) on their multi-center data and found limited cross-cohort generalizability (Table S1), which motivated building a dedicated pipeline.
Methodology in Plain English
Data. Four datasets from three countries were used. Model development used the CogTrAiL-RBD randomized controlled trial cohort (Germany): 78 individuals (55 iRBD, 23 HC), 524 nights. The blinded local test set came from an ongoing iRBD screening effort at the same site: 31 participants (19 iRBD, 12 HC). Two external cohorts were used: the Oxford Discovery (OPDC) cohort of 103 individuals (70 iRBD, 8 PD+RBD, 25 HC) yielding 113 actigraphy samples from longitudinal follow-up, and the PACE cohort from Denmark with 31 individuals (13 iRBD, 10 PD+RBD, 5 PD–RBD, 3 HC) in the intro description, producing 57 samples. All recordings used the Axivity AX6 at a 100 Hz sample rate and ±8 g dynamic range, typically spanning 6–7 consecutive nights and totaling more than 1,809 nights across all cohorts.
Preprocessing. Raw signals are resampled to correct internal clock drift, bandpass filtered, screened for non-wear episodes, auto-calibrated using the fact that the acceleration vector magnitude should approximate standard gravity during rest (following van Hees et al., 2014), and segmented into sleep and wake periods using the HDCZA algorithm (van Hees et al., 2018). This replaces subjective sleep diaries and manual calibration.
Feature extraction. From the detected sleep bouts, the team computed interpretable motion descriptors — intensity, periodicity, spectral separation of rapid versus slow movements, complexity, fragmentation, overall activity, and clustering of movement bursts across the night. Local features per activity bout were aggregated into global per-night descriptors, with a sleep expert involved in the design.
Modeling. Each night's global features are mapped to an RBD probability by a boosted decision tree classifier (XGBoost). Nightly scores are then combined into a patient-level risk score using a custom aggregation function that mixes mean-probability thresholding with majority voting, and a threshold converts this into a binary prediction. Feature selection and hyperparameter tuning were done inside a nested cross-validation framework to avoid biased performance estimates; the final model was trained on the full development cohort and evaluated once on the blinded holdout.
Evaluation. Performance was reported at both night and patient level using AUROC, F1, and balanced accuracy, with 95% confidence intervals from inter-fold variability for cross-validation and bootstrap resampling (n = 2000, stratified by class) for single-model held-out evaluations. Robustness was probed with leave-one-dataset-out cross-validation, and an exploratory cross-device evaluation used a small publicly available single-night cohort of 28 participants (3 RBD, 25 non-RBD) under lower prevalence conditions. The detailed results of that cross-device evaluation are not included in the provided text.
Why This Matters
This work matters because it moves actigraphy-based RBD detection from promising single-center experiments toward a reproducible, openly available tool that others can run, test, and improve on their own data. The quantitative preprocessing validation — calibration error, filter performance, sleep-detection agreement — is unusual for this literature and makes the pipeline auditable rather than a black box. The multi-cohort design, especially the blinded holdout and two external cohorts, provides a much more honest picture of real-world generalization than internal cross-validation alone.
Real-world applications:
- Large-scale pre-screening for prodromal α-synucleinopathies in research cohorts and population studies, where vPSG is impractical.
- Enrichment of clinical trials by identifying likely iRBD participants from low-cost wearable recordings.
- Low-burden longitudinal monitoring in memory and movement disorder clinics, using devices patients already tolerate.
- Reuse of the preprocessing module as a standalone, general-purpose actigraphy analysis tool independent of RBD detection.
Industry relevance: For wearable and digital-health companies, the paper highlights the value of calibrated, full-resolution tri-axial data and standardized, device-agnostic pipelines over ad hoc per-study analysis. The focus on automated sleep segmentation without diaries, a pretrained model for immediate use, and computational efficiency makes the approach deployable in screening products and remote-monitoring platforms — and the open-source release invites adoption, independent validation, and cumulative refinement across commercial and academic settings.
Future Directions
- Broader and lower-prevalence validation. The authors note that scalable deployment requires validation beyond the cohorts tested here; the exploratory cross-device evaluation used only 28 participants (3 RBD, 25 non-RBD), and its reported results are not detailed in the provided text.
- Pooled multi-center learning. The paper argues that progress depends on moving beyond parallel model development toward integrative benchmarking and pooled learning across centers, and presents the pooled multi-center pre-trained model as a step in that direction.
- Community-driven benchmarking and independent replication. Because ActiTect is open-source and easy to use, the stated hope is that others will independently validate it, compare it systematically against alternatives, and contribute improvements.
- Handling clinical heterogeneity and class imbalance. Balanced accuracy dropped for PD+RBD tasks on both external cohorts (0.7765 on PACE and 0.81309 on OPDC) despite preserved AUROC, pointing to threshold calibration in more heterogeneous and imbalanced populations as an open problem.
Target Audience
This paper is most valuable to clinical machine learning researchers and biomedical engineers working on wearable sensing and neurodegenerative disease biomarkers; clinicians and sleep specialists interested in scalable RBD screening; and data scientists at digital-health or wearable companies who need a concrete example of cross-device harmonization, automated sleep segmentation, and rigorous external validation. Readers with a basic grasp of classification metrics and cross-validation will be able to follow it, and those looking to build or deploy actigraphy-based screening tools will find the pipeline design and validation strategy directly actionable.
Authors’ abstract
Isolated rapid eye movement sleep behavior disorder (iRBD) is a major prodromal marker of $α$-synucleinopathies, often preceding the clinical onset of Parkinson's disease, dementia with Lewy bodies, or multiple system atrophy. While wrist-worn actimeters hold significant potential for detecting RBD in large-scale screening efforts by capturing abnormal nocturnal movements, they become inoperable without a reliable and efficient analysis pipeline. This study presents ActiTect, a fully automated, open-source machine learning tool to identify RBD from actigraphy recordings. To ensure generalizability across heterogeneous acquisition settings, our pipeline includes robust preprocessing and automated sleep-wake detection to harmonize multi-device data and extract physiologically interpretable motion features characterizing activity patterns. Model development was conducted on a cohort of 78 individuals, yielding strong discrimination under nested cross-validation (AUROC = 0.95). Generalization was confirmed on a blinded local test set (n = 31, AUROC = 0.86) and on two independent external cohorts (n = 113, AUROC = 0.84; n = 57, AUROC = 0.94). To assess real-world robustness, leave-one-dataset-out cross-validation across the internal and external cohorts demonstrated consistent performance (AUROC range = 0.84-0.89). A complementary stability analysis showed that key predictive features remained reproducible across datasets, supporting the final pooled multi-center model as a robust pre-trained resource for broader deployment. By being open-source and easy to use, our tool promotes widespread adoption and facilitates independent validation and collaborative improvements, thereby advancing the field toward a unified and generalizable RBD detection model using wearable devices.