Research
Enhancing Time Series Classification with Diversity-Driven Neural Network Ensembles
Overview Research area: Machine learning — specifically deep learning for Time Series Classification (TSC) and neural network ensemble methods. Technical level: Intermediate. The paper assumes familia

- arXiv
- 2602.07579
- Published
- 2026-02-07
- Authors
- Javidan Abdullayev, Maxime Devanne, Cyril Meyer, Ali Ismail-Fawaz, Jonathan Weber, Germain Forestier
AI summary
Overview
Research area: Machine learning — specifically deep learning for Time Series Classification (TSC) and neural network ensemble methods.
Technical level: Intermediate. The paper assumes familiarity with convolutional neural networks, loss functions, and ensemble learning, but its central idea (make ensemble members learn different features) is explained in accessible terms.
Scope: The paper proposes a diversity-driven ensemble framework built on the LITE architecture, applies a feature orthogonality loss during sequential training of ensemble members, and evaluates it on 128 UCR archive datasets against ensembles of identical architectures and against several state-of-the-art methods.
What This Paper Is About
State-of-the-art deep learning ensembles for time series classification, such as InceptionTime, H-InceptionTime and LITETime, train multiple copies of the same architecture with different random initializations and simply average their predictions. Because nothing forces those copies to learn different things, they often converge to similar feature representations, making part of the ensemble redundant. This paper's goal is to explicitly push ensemble members to learn complementary features using a feature orthogonality loss, so that strong accuracy can be reached with fewer models.
Key Contributions
- A diversity-driven ensemble framework that explicitly promotes feature diversity in TSC ensembles, based on a feature orthogonality loss applied to the learned feature representations rather than to the convolutional filters.
- An empirical validation across 128 UCR datasets showing SOTA-level performance with fewer models than the standard LITETime ensemble, improving efficiency.
- A demonstration that filter-level orthogonality alone does not produce meaningful diversity: filters become orthogonal while feature maps remain similar.
- Quantitative and qualitative diversity analysis using Fréchet Inception Distance (FID) scores and t-SNE visualizations of learned convolutional filters.
Main Findings
-
Decorrelation beats a larger base ensemble on the UCR archive: The 4-model decorrelated ensemble (Deco-LITETime-4) achieves the highest mean accuracy among all configurations evaluated, surpassing the state-of-the-art LITETime-5, which uses five LITE classifiers. The p-value between Deco-LITETime-4 and LITETime-5 indicates the two are not statistically different, while the p-value between LITETime-4 and LITETime-5 is lower than 5%, indicating a statistically significant difference.
-
Fewer models, comparable results: The 3-model decorrelated ensemble (Deco-LITETime-3) performs comparably to LITETime-5, with the difference not statistically significant, positioning it as a viable alternative to the SOTA with two fewer models.
-
Gains grow with ensemble size: Benefits are less evident in small ensembles. Deco-LITETime-2 shows only marginal improvement over its base counterpart LITETime-2 and does not fully bridge the gap to LITETime-5. The paper attributes this to base models already exhibiting some inherent diversity in small ensembles, with overlap increasing as the ensemble grows. The paper also separately reports that Deco-LITETime-2 achieves approximately 3.6% higher test accuracy than LITETime-5.
-
Third place against state-of-the-art methods, with far fewer parameters: In the multi-comparison against SOTA approaches, the proposed method achieves an average accuracy of 0.8496, ranking third. Only MultiROCKET significantly outperforms it. Deco-LITETime-4 contains less than 10% of the parameters of a single Inception-based classifier.
-
Feature diversity is measurably higher: FID comparison across 128 UCR datasets shows the decorrelated model produces higher FID scores in 90 datasets, often with a larger margin, while the base model gives higher scores in 32 datasets; in 6 datasets no difference was observed. The p-value between the two sets of FID scores is 0.0.
-
Filter-level orthogonality was rejected as insufficient: Experiments showed that applying orthogonality loss to convolutional filters tended to shift discriminative patterns rather than create diverse representations — filters became orthogonal but feature maps stayed similar.
-
Base models learn redundant filters: t-SNE visualization of the final-layer filters of five base models (32 filters, filter dimension 32 × 20) shows a highly similar distribution forming two clusters that each contain filters from all base models. With the first base model as reference plus four decorrelated models, diversity between base and decorrelated filters is evident, and the decorrelated models also differ from each other. The fifth decorrelated model (Deco_5) distributes similarly to Deco_2, suggesting it fails to add further diversity because the most useful diverse features were already captured.
-
Illustrative single-dataset result: On the BirdChicken dataset, base LITE models with different initializations produce highly similar feature maps, whereas the decorrelated model learns different features, enabling the decorrelated ensemble to reach 100% test accuracy versus 90% for the base ensemble, and exceeding the maximum 95% test accuracy of InceptionTime's five-model ensemble.
Methodology in Plain English
The researchers start from LITE, a lightweight inception-style convolutional model, and replace its usual "train five copies independently" recipe with a sequential, coordinated procedure. The first model in the ensemble is trained normally with cross-entropy loss only. Every subsequent model is trained with an extra penalty — a feature orthogonality loss — that measures the cosine similarity between its own feature outputs and the feature outputs of all previously trained models. Minimizing that penalty pushes the new model toward feature representations that are unlike those already in the ensemble. The penalty is averaged over all predecessors and added to the classification loss, with a weight of α = 0.5, giving equal importance to accuracy and diversity.
A deliberate choice is to apply the penalty only to the final convolutional layer, not to early layers. The reasoning is that lower layers learn generic features and constraining them can hurt performance, while upper layers hold task-specific representations where redundancy is the real problem. The authors also explain why they did not use misclassification-based diversity methods such as Deep Negative Correlation Classification: the LITE model reaches 100% training accuracy on most UCR datasets, leaving little room for misclassification-driven diversity.
Comparison is done with standard, identical training settings: Adam optimizer, initial learning rate 0.001, reducing factor 0.5, patience 50, 1500 epochs, batch size 64, batch size 64 inputs on z-normalized time series using the original UCR train/test splits. Decorrelated models reuse the same initialization seeds as their base counterparts so that improvements can be attributed to the orthogonality loss rather than to randomness. Each ensemble configuration is trained five times and averaged. Evaluation uses accuracy across 128 datasets with the Multi-Comparison Matrix (Mean Accuracy, Mean Difference, Win/Tie/Loss) and the Wilcoxon signed-rank test at p < 0.05. Diversity is measured with FID scores and inspected with t-SNE projections of filters, with Dynamic Time Warping used to measure filter similarity.
Why This Matters
Impact on research: The paper challenges the common assumption that simply training more identically configured models will keep improving an ensemble. It shows that diversity is a designable property rather than a byproduct of random initialization, and that explicitly penalizing redundant features lets a smaller ensemble match or exceed a larger one. It also contributes a negative result about filter-level orthogonality, redirecting attention to feature-level constraints.
- Healthcare: Time series classification is listed among the application domains, where patient monitoring signals must be classified accurately and cheaply.
- Human activity recognition: The paper lists this as a core application area for TSC, where wearable sensor streams need reliable classification.
- Social security: Named as an application domain in the introduction.
- Remote sensing: Named as an application domain, where temporal satellite or sensor sequences are classified.
Industry relevance: The method reduces the number of models needed for state-of-the-art accuracy, and Deco-LITETime-4 uses less than 10% of the parameters of a single Inception-based classifier. Fewer models mean lower training and inference cost, which matters for deployment on constrained hardware and for scaling to large dataset collections. The code is publicly released at https://github.com/MSD-IRIMAS/decorrelated-learning.
Future Directions
- Refining the balance between the classification loss and the diversity loss, and optimizing that trade-off dynamically across different datasets. The paper notes that diversity loss against the base model appears easier to optimize than diversity loss against previously trained decorrelated models, and that per-term weighting could improve results.
- Extending the framework to multivariate time series classification, which the current work does not cover.
- Investigating alternative diversity-promoting strategies beyond feature orthogonality.
- Improving computational efficiency, potentially through parallel training strategies to enhance scalability, since the current procedure trains models sequentially.
Target Audience
Researchers and practitioners working on time series classification, deep learning ensembles, and efficient model design. It is most useful to readers who already understand convolutional architectures and loss functions and want to know how to make an existing ensemble pipeline more diverse and more efficient. Readers interested in the practical trade-off between accuracy and model count, or in reproducible benchmarks on the UCR archive, will also benefit. The paper reports details not covered here, such as exact FID values and exact p-values for most comparisons, which readers should consult in the original figures and tables.
Authors’ abstract
Ensemble methods have played a crucial role in achieving state-of-the-art (SOTA) performance across various machine learning tasks by leveraging the diversity of features learned by individual models. In Time Series Classification (TSC), ensembles have proven highly effective whether based on neural networks (NNs) or traditional methods like HIVE-COTE. However most existing NN-based ensemble methods for TSC train multiple models with identical architectures and configurations. These ensembles aggregate predictions without explicitly promoting diversity which often leads to redundant feature representations and limits the benefits of ensembling. In this work, we introduce a diversity-driven ensemble learning framework that explicitly encourages feature diversity among neural network ensemble members. Our approach employs a decorrelated learning strategy using a feature orthogonality loss applied directly to the learned feature representations. This ensures that each model in the ensemble captures complementary rather than redundant information. We evaluate our framework on 128 datasets from the UCR archive and show that it achieves SOTA performance with fewer models. This makes our method both efficient and scalable compared to conventional NN-based ensemble approaches.