Research
Group and Exclusive Sparse Regularization-based Continual Learning of CNNs
Overview Research area: Continual learning (also called lifelong or sequential learning) for convolutional neural networks, specifically regularization-based methods that fight catastrophic forgetting

- arXiv
- 2601.03658
- Published
- 2026-01-07
- Authors
- Basile Tousside, Janis Mohr, Jörg Frochte
AI summary
Overview
Research area: Continual learning (also called lifelong or sequential learning) for convolutional neural networks, specifically regularization-based methods that fight catastrophic forgetting.
Technical level: Advanced. The paper derives a composite regularized loss, a proximal gradient descent update rule per convolution filter, and a filter importance measure, so comfort with optimization and CNN internals helps.
One-sentence scope: The paper proposes Group and Exclusive Sparsity based Continual Learning (GESCL), a fixed-capacity CNN method that pairs a stability regularizer (protecting important filters) with a plasticity regularizer (sparsifying unimportant filters) and evaluates it on Split SVHN, Split CIFAR-10/100 and ImageNet-50.
What This Paper Is About
A CNN trained on a sequence of tasks tends to lose accuracy on earlier tasks once new data arrives, a failure mode known as catastrophic forgetting. The core difficulty is the stability-plasticity dilemma: the network must stay stable enough to keep solving old tasks while remaining plastic enough to learn new ones. The authors' goal is to resolve this dilemma inside a fixed-capacity CNN, without growing the network or storing data from past tasks.
Key Contributions
- A plasticity regularizer combining exclusive and group sparsity. The method constrains convolution kernels with a group sparsity term (which deactivates whole filters and promotes feature sharing) and an exclusive sparsity term (which pushes filters to learn disjoint, discriminating features). The balance between the two is layer-dependent, with group sparsity weighted more in lower layers and exclusive sparsity dominating in higher layers.
- A stability regularizer applied on top of the sparsity term. Filter parameters learned up to task t-1 are penalized from deviating during training on task t, but only for the subset of filters identified as important for the tasks seen so far.
- An adaptive, post-activation-based filter importance measure. Instead of treating all important filters as equally important, importance is computed as the average standard deviation of a filter's ReLU post-activation values over the training samples of a task, then accumulated across tasks with a balancing hyperparameter.
- A proximal gradient descent solver with closed-form per-filter updates. Because convolution filters do not overlap, the regularized objective decomposes and the proximal step can be applied independently to each filter, with explicit update rules for important filters and for unimportant filters, followed by reinitialization of unimportant filters and zeroing of the corresponding downstream channels.
Main Findings
- Strongest retention on the initial task. Across Split SVHN, Split CIFAR-10/100 and ImageNet-50, GESCL showed the least degradation of first-task accuracy as further tasks were learned (Figure 2). On the CIFAR-10/100 dataset, GESCL also showed strong per-task retention as new tasks were learned (Figure 3).
- HAT is a close competitor on small benchmarks. HAT achieved high performance retention especially on the small SVHN dataset, while EWC suffered severe degradation, particularly on ImageNet-50 and CIFAR-10/100, which the authors attribute to EWC's limitation when the number of tasks grows large.
- Best average accuracy. Using the average accuracy metric (mean of accuracies on tasks 1 through t), GESCL outperformed the baselines during the continual learning runs (Figure 4).
- SI catches up late on ImageNet-50. SI reached the best performance on the last task of ImageNet-50 despite being clearly outperformed by GESCL on earlier tasks. AGS was the strongest competitor in average accuracy, while HAT surpassed AGS on SVHN and ImageNet in terms of avoiding forgetting.
- Every component of GESCL contributes. In the ablation study on CIFAR-10/100 and ImageNet-50, the full model with the stability regularizer, plasticity regularizer and filter importance balance all active reached average accuracy (A) of 74.5% with average forgetting (F) of 0.0008 on CIFAR-10/100, and 77.1% with F of 0.00007 on ImageNet-50.
- Removing the importance balance hurts. Deactivating the filter importance balance gave 71.6% A / 0.013 F on CIFAR-10/100 and 75.10% A / 0.068 F on ImageNet-50.
- Removing the stability regularizer is the largest single loss. Without the stability regularizer, results dropped to 41.9% A / 0.44 F on CIFAR-10/100 and 57.38% A / 0.349 F on ImageNet-50.
- Removing the plasticity regularizer also degrades performance. Without plasticity regularization, results were 69.12% A / 0.013 F on CIFAR-10/100 and 73.59% A / 0.018 F on ImageNet-50.
- Reported comparison values for the baselines. The paper presents the baseline comparisons in Figures 2, 3 and 4; numerical accuracy and forgetting values for EWC, SI, MAS, HAT and AGS are not reported in the text.
- Parameter and compute efficiency. The authors state qualitatively that GESCL uses significantly fewer parameters and less computation than approaches that dynamically expand the network or memorize past task data. No specific parameter counts or compute measurements are reported.
Methodology in Plain English
The authors start from a standard CNN classifier and train it one task at a time, discarding each task's training data after it has been learned.
After training on a task, they look at how strongly each convolution filter responds to the training images, measured by the standard deviation of the filter's post-activation values (the intuition being that filters producing weak activations are not detecting useful features). Each filter gets an importance score, and scores are carried forward across tasks using a weighted combination of the previous accumulated score and the new one.
Filters above the importance threshold form the "important" set and are constrained by the stability regularizer so that they stay close to the values they had at the end of the previous task. Filters with a zero importance score form the "unimportant" set and are instead pushed toward zero by the plasticity regularizer, which applies both a group sparsity term (shrinking whole filters) and an exclusive sparsity term (making filters learn as distinct features as possible). The relative weighting of the two sparsity terms changes with depth: group sparsity is weighted more in lower layers, exclusive sparsity in higher layers.
Unimportant filters whose kernels have been zeroed out are randomly reinitialized so they can be reused for future tasks, and the matching channel in the next layer is zeroed so it does not corrupt inference.
Because the objective mixes a smooth cross-entropy term with non-smooth regularization terms, training uses proximal gradient descent: a normal gradient step on the cross-entropy loss, followed by a corrective proximal step. Since filters do not overlap, this corrective step is applied independently per filter, yielding simple closed-form rules that interpolate between the old filter weights and the new ones for important filters, and shrink toward zero for unimportant filters.
Experiments use Split SVHN (10 classes grouped into 5 tasks), Split CIFAR-10/100 (the 10 CIFAR-10 classes as the first task plus CIFAR-100 split into 10 tasks, giving 11 tasks), and ImageNet-50 (a generated subset of ImageNet with 50 classes grouped into pairs to form 25 tasks, with 1600 training images and 200 validation images per class). The SVHN and CIFAR networks use 3 blocks of 3x3 convolutions with 32, 64 and 128 filters, ReLU activations and 2x2 max-pooling; the ImageNet-50 network uses two blocks of 2x2 convolution with 64 filters, ReLU and 2x2 max-pooling. All experiments use a multi-headed network. Baselines are EWC and SI as reference methods plus MAS, HAT and AGS, with grid search used to select hyper-parameters fairly for each approach.
Why This Matters
The work targets a practical constraint that shows up whenever a model cannot revisit its past data: retaining old skills without storing old samples or inflating the model. Its distinctive angle is that it treats stability and plasticity as two regularizers acting on the same fixed set of filters, rather than growing the network or keeping a memory buffer, and it treats filter importance as graded rather than binary-and-uniform.
Real-world applications:
- Privacy-constrained systems, where regulations or policies require deleting user data after a period, so models must adapt to new data without revisiting deleted samples.
- Online and streaming learning, where data for new tasks arrives on the fly and the model must update without a full retraining pass over history.
- Edge and embedded vision devices, where fixed memory budgets make dynamically expanding architectures or replay buffers impractical.
- Industrial inspection and medical imaging pipelines, where new classes of defects, lesions or equipment variants appear over time and a deployed model must absorb them without losing accuracy on the original cases.
Industry relevance: Because GESCL keeps network capacity fixed and avoids storing past data, it fits deployment scenarios where memory, storage and compute are constrained, and where replaying old data is legally or logistically impossible. The authors also position it against replay-based and architecture-expansion methods in terms of parameter and computation cost, though no quantitative efficiency measurements are provided.
Future Directions
- Quantifying the efficiency claim. The paper asserts lower parameter and computation cost than expansion-based and memory-based approaches, but no measured parameter counts, memory footprints or training times are reported, leaving a direct efficiency comparison as an open step.
- Scalability of the stability-plasticity balance. The authors report that with only the stability term the CNN becomes inefficient at learning future tasks on a 25-task benchmark; how the regularizer strengths and the filter importance balance should be set as the number of tasks grows further is not established.
- Head-to-head analysis of competitor strengths. SI reaching the best final-task accuracy on ImageNet-50 while losing on earlier tasks, and HAT beating AGS on forgetting but losing to it on average accuracy, are noted but not explained; understanding these trade-offs could inform hybrid designs.
- Beyond image classification. The method is built around convolution filters and ReLU post-activations, and the authors explicitly frame feature sharing and discrimination in terms of image classification, so extension to other architectures and modalities is unaddressed. The paper does not state explicit future work in its conclusion.
Target Audience
Researchers and graduate students working on continual learning, catastrophic forgetting and network regularization; practitioners who need to update deployed vision models sequentially under fixed memory or data-retention constraints; and readers interested in structured sparsity, filter pruning and proximal gradient optimization applied to training objectives. Readers without a background in CNN internals or convex optimization will find the method section demanding, while the experimental results and ablation table are accessible on their own.
Authors’ abstract
We present a regularization-based approach for continual learning (CL) of fixed capacity convolutional neural networks (CNN) that does not suffer from the problem of catastrophic forgetting when learning multiple tasks sequentially. This method referred to as Group and Exclusive Sparsity based Continual Learning (GESCL) avoids forgetting of previous tasks by ensuring the stability of the CNN via a stability regularization term, which prevents filters detected as important for past tasks to deviate too much when learning a new task. On top of that, GESCL makes the network plastic via a plasticity regularization term that leverage the over-parameterization of CNNs to efficiently sparsify the network and tunes unimportant filters making them relevant for future tasks. Doing so, GESCL deals with significantly less parameters and computation compared to CL approaches that either dynamically expand the network or memorize past tasks' data. Experiments on popular CL vision benchmarks show that GESCL leads to significant improvements over state-of-the-art method in terms of overall CL performance, as measured by classification accuracy as well as in terms of avoiding catastrophic forgetting.