Research
ProtoTSNet: Interpretable Multivariate Time Series Classification With Prototypical Parts
Overview Research area: Interpretable machine learning / explainable AI applied to multivariate time series classification (MTSC). Technical level: Advanced. The paper assumes familiarity with convolu

- arXiv
- 2511.02152
- Published
- 2025-11-04
- Authors
- Bartłomiej Małkus, Szymon Bobek, Grzegorz J. Nalepa
AI summary
Overview
Research area: Interpretable machine learning / explainable AI applied to multivariate time series classification (MTSC).
Technical level: Advanced. The paper assumes familiarity with convolutional neural networks, prototype-based explanation methods (ProtoPNet), autoencoders, and benchmark evaluation with critical difference diagrams.
Scope: The paper introduces ProtoTSNet, an ante-hoc explainable neural architecture for multivariate time series, and benchmarks it against explainable and non-explainable baselines on 30 datasets from the UEA archive.
What This Paper Is About
Time series classification models in high-stakes domains such as healthcare and industry are often accurate but opaque, and post-hoc explanations generated after training can misrepresent what a black-box model actually does. The authors adapt the prototype-based ProtoPNet idea from computer vision to multivariate time series, where the model classifies by matching input subsequences against learned prototypical parts drawn from the training data. The goal is a model whose explanations are built into its structure while still remaining competitive with non-explainable and post-hoc explainable methods.
Key Contributions
-
A modified convolutional encoder for time series prototypes. ProtoTSNet uses input feature masking combined with group convolutions, where each of
lencoder groups receives a different randomly selected subset of input features (controlled by a reception parameterr). This prevents information from informative and uninformative dimensions from being mixed during prototype selection, a problem that does not arise in image data where all channels are typically meaningful. -
Built-in, quantifiable feature importance. A
1 x 1convolution layer mixes the group encoder outputs, and the importance of each input feature is computed directly from the masking pattern and the mixing weights (I_m = sum_j |sum_i delta_im * w_ij|). This yields global feature importance scores rather than prototype-specific ones. -
A staged training procedure with encoder pretraining. The encoder is first trained as part of an autoencoder with MSE loss, then the full model is trained in repeating cycles of warm epochs, joint epochs, prototype projection onto the nearest latent training patch, and last-layer optimization, using a composite loss with clustering, separation, and L1 sparsity terms.
-
Reproducible benchmarking with re-run baselines. The authors evaluate on 30 UEA multivariate datasets, rerun competing methods to verify their reported results, and publish source code at https://github.com/bmalkus/ProtoTSNet.
Main Findings
-
Best among ante-hoc explainable methods. ProtoTSNet achieves an average rank of 3.90 across the UEA benchmarks, the best result among ante-hoc explainable methods (ProtoTSNet, PETSC, and the shapelet-based method).
-
Non-explainable methods still lead overall. ROCKET achieves the best average rank of 2.73, followed by the post-hoc explainable method LITEMVTime at 3.03. The authors attribute this hierarchy to the accuracy advantage of methods not constrained by ante-hoc explainability requirements.
-
Statistical analysis supports competitiveness. Using a critical difference test with Nemenyi post-hoc analysis at alpha = 0.05 (diagram generated with the Orange data mining library), ProtoTSNet shows comparable performance to non-explainable methods and post-hoc explainable methods.
-
Dataset characteristics vary widely. The 30 UEA datasets span 2 to 39 classes, 2 to 1345 features, and sequence lengths from 8 to 17984.
-
Dataset-specific behavior. ProtoTSNet performs strongly on DuckDuckGeese and HandMovementDirection, where convolutional filtering and temporal pattern extraction help. Shapelet-based and post-hoc methods perform better on EthanolConcentration and Handwriting, suggesting those datasets favor global statistical features or fine-grained local patterns rather than the intermediate-length prototypical parts.
-
Synthetic validation of interpretability. On a synthetic dataset with three features (two significant, one not), four classes, and class-discriminative patterns in only the initial 40 time steps (the following 60 randomized), the model reached 100% classification accuracy. It correctly captured the significant fragments as prototypes and dimmed the insignificant feature in the computed feature importance.
-
Qualitative prototype inspection. On the Libras dataset (15 classes, 24 instances per class), the authors fixed the number of prototypes to 2 per class, down from 10 used in the UEA dataset experiments, and varied the last-layer L1 coefficient for qualitative analysis.
-
Baseline reproduction was imperfect. The authors' experiments did not replicate the reported performance of XCM and MTEX-CNN, and LAXCAT was excluded from comparison due to unclear preprocessing steps and no available source code.
Methodology in Plain English
The model starts by making several copies of each input time series, each copy with a different random subset of its features switched off. The size of that subset is set by a reception parameter r, and the number of copies equals the number of encoder groups l (typically 32). Each copy is then passed through its own independent convolutional group, so features from different subsets never mix inside the encoder.
The outputs of these groups are combined by a small mixing layer, and the weights of that layer, together with the recorded masking pattern, directly reveal how much each original input feature matters to the model. This is how feature importance is obtained without a separate explanation step.
In the resulting latent space, the model holds a set of learned prototypes, each a short subsequence spanning all encoded features over a fixed number of time steps. For every prototype, the model scans all subsequences of the input and computes an L2 distance, converting the closest match into a similarity score. The highest similarity score per prototype becomes that prototype's activation, and a final dense layer combines those activations into class predictions. Because time order is preserved, each prototype can be mapped back to the actual segment of the input it resembles, which is what makes the explanation human-readable.
Training happens in two phases. First, the encoder is trained inside an autoencoder to reconstruct the input. Second, the full model is trained with prototypes as learnable parameters, initialized so that a prototype's own class gets weight 1 and other classes get weight -0.5, with biases disabled. Training then alternates through warm epochs (only prototypes and the mixing layer train), joint epochs (everything trains), projection of each prototype onto the nearest real latent patch from the training data, and last-layer-only optimization, repeated in cycles.
Evaluation uses a synthetic dataset for interpretability checks and 30 UEA datasets for comparison against ante-hoc explainable methods (PETSC, a shapelet-based classifier built with the sktime ShapeletTransformClassifier v0.27.0 using RotationForestClassifier), post-hoc explainable methods (XCM, MTEX-CNN, LITEMVTime), and non-explainable methods (ROCKET, TapNet). Hyperparameters were grid-searched with 5-fold cross-validation over r in {0.25, 0.5, 0.75, 0.9} and prototype length L at 1%, 10%, 25%, 50%, and 100% of the series length. Training comprised 50 encoder pretraining epochs plus roughly 360 epochs of main training across four cycles.
Why This Matters
Ante-hoc explainability matters because explanations produced after the fact can be inaccurate or create misconceptions about how a model actually behaves, a problem that is unacceptable when decisions carry significant consequences in regulated settings such as those governed by GDPR and the AI Act. The paper also addresses two gaps identified in surveys of explainable time series classification: most work focuses on univariate series, and much of it lacks reproducible code.
Real-world applications mentioned or implied by the paper:
- Healthcare and medicine, where classification decisions have critical consequences.
- Industrial monitoring, where time series data is a primary modality.
- Finance, where model transparency is increasingly required.
- Human rehabilitation exercise evaluation, the demonstrated application domain of the LITEMVTime baseline.
Industry relevance: Because ProtoTSNet produces feature importance scores and prototypes that map back to interpretable time segments, domain experts can inspect which time ranges and which sensors or variables drove a prediction without needing to run a separate explanation tool. The authors position the method as scalable and competitive with non-explainable and post-hoc explainable approaches, which matters for deployment where transparency requirements conflict with the accuracy of black-box models.
Future Directions
-
Improving accuracy relative to non-explainable methods. ProtoTSNet ranks behind ROCKET and LITEMVTime in average rank, so narrowing that gap while preserving ante-hoc interpretability is an open problem.
-
Reception parameter trade-off. The authors state that larger
rvalues risk overestimating feature importance through increased feature interaction, while smallerrvalues preserve feature separation but may miss contextual relationships. How to chooseroptimally remains an open question. -
Multivariate explainability beyond global scores. The current feature importance calculation produces global scores across all prototypes rather than prototype-specific values; extending it to per-prototype importance is a natural next step.
-
Comparisons that could not be completed. LAXCAT was excluded for lack of preprocessing detail and source code, and XCM and MTEX-CNN results could not be replicated. Restoring these comparisons under a common protocol would strengthen the picture of where ante-hoc methods stand.
Target Audience
Researchers working on explainable AI for time series, practitioners in regulated domains such as healthcare, industry, and finance who need transparent classification models, and machine learning engineers interested in prototype-based architectures and reproducible benchmark evaluation. Readers looking for an introductory treatment of time series classification will find the paper assumes substantial prior knowledge of convolutional architectures and explainability literature.
Authors’ abstract
Time series data is one of the most popular data modalities in critical domains such as industry and medicine. The demand for algorithms that not only exhibit high accuracy but also offer interpretability is crucial in such fields, as decisions made there bear significant consequences. In this paper, we present ProtoTSNet, a novel approach to interpretable classification of multivariate time series data, through substantial enhancements to the ProtoPNet architecture. Our method is tailored to overcome the unique challenges of time series analysis, including capturing dynamic patterns and handling varying feature significance. Central to our innovation is a modified convolutional encoder utilizing group convolutions, pre-trainable as part of an autoencoder and designed to preserve and quantify feature importance. We evaluated our model on 30 multivariate time series datasets from the UEA archive, comparing our approach with existing explainable methods as well as non-explainable baselines. Through comprehensive evaluation and ablation studies, we demonstrate that our approach achieves the best performance among ante-hoc explainable methods while maintaining competitive performance with non-explainable and post-hoc explainable approaches, providing interpretable results accessible to domain experts.