Skip to content
AI.info

Research

C-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression

Overview Research area: Model compression for computer vision — specifically one-shot structured pruning of deep neural networks, guided by explainable AI (XAI) and causal inference. Technical level:

C-SWAP: Explainability-Aware Structured Pruning for Efficient Neural Networks Compression
arXiv
2510.18636
Published
2025-10-21
Authors
Baptiste Bauvin, Loïc Baret, Ola Ahmad

AI summary

Overview

Research area: Model compression for computer vision — specifically one-shot structured pruning of deep neural networks, guided by explainable AI (XAI) and causal inference.

Technical level: Intermediate. The paper assumes familiarity with structured pruning terminology (channels, filters, neurons), attribution methods such as Integrated Gradients, DeepLIFT, and LRP, and basic hypothesis testing. It does not assume prior expertise in causal inference, which is introduced from the ground up.

Scope: The paper proposes C-SWAP, an explainability-aware one-shot structured pruning framework that identifies neurons as critical, neutral, or detrimental through a multi-class causal effect criterion, and progressively removes the non-critical ones without any fine-tuning.

What This Paper Is About

Structured pruning can make deep networks smaller and faster by deleting entire neurons, channels, or layers, but the standard approach requires repeated cycles of pruning and retraining, which is costly and unstable. One-shot post-training pruning avoids that cost but typically drops accuracy sharply at high pruning ratios, because single global rankings of neuron importance become unreliable once many units have already been removed. The paper's goal is to prune aggressively in one shot, with no fine-tuning, by using causal explanations to decide which neurons are safe to delete as the network is progressively modified.

Key Contributions

  1. A multi-class causal explanation criterion. The authors generalize a class-specific causal effect measure into a global causal effect ξ_n computed over samples from all C classes, and use it — combined with a per-class paired t-test and a voting strategy — to label each neuron as critical, neutral, or detrimental.

  2. The C-SWAP algorithm. A causal-guided pruning algorithm for deep and complex architectures that interleaves analysis and pruning, processing the network from the output layer back toward the input. It avoids global rankings, running in O(n) rather than the O(m × n) of a greedy prune-one-and-rerank procedure, and requires no fine-tuning.

  3. Empirical superiority across CNNs and a vision transformer. Experiments on four architectures and two classification datasets show C-SWAP outperforming all considered baselines, including magnitude pruning (OMP), random pruning, several attribution-based methods, and an adapted version of ACDC-style mechanistic pruning.

  4. Extension to semantic segmentation. By substituting mean Intersection-over-Union (mIoU) as the scoring function, C-SWAP is applied to DDRNet 23S on the Cityscapes dataset, demonstrating that the framework is not limited to classification.

Main Findings

  • C-SWAP achieves the best SAUCE scores on every architecture tested. Under the newly proposed Sparsity AUC Estimator (SAUCE), C-SWAP records 72.1 ± 0.39 on ResNet-18/CIFAR10, 80.3 ± 0.28 on ResNet-50/ImageNet, 49.4 ± 1.15 on MobileNetV2/ImageNet, and 68.1 ± 1.67 on ViT/ImageNet, for an average of 67.5. The next-best averages are Internal Influence and iLRP at 40.3.

  • No fine-tuning is needed at moderate pruning. On ResNet-18 with CIFAR10, C-SWAP removes up to 50% of parameters without performance loss, surpassing all other methods in the comparison.

  • Baselines degrade quickly. Random pruning averages 27.4 SAUCE, OMP averages 29.1, and Adapted Mechanistic Pruning (AMP) averages 33.6 — despite also using a progressive strategy, AMP fails to preserve essential structures. DeepLIFT, Conductance, and Integrated Gradients cluster around 37.9 to 39.9 on average.

  • Progressive pruning matters. C-BP, an ablation of C-SWAP that removes the progressive component and simply prunes the least important neurons by the causal criterion, averages 39.7 SAUCE and varies considerably across seeds (for example, 42.3 ± 9.73 on ResNet-50/ImageNet), indicating ranking instability.

  • Two results are unavailable due to implementation incompatibilities. iLRP could not be run on MobileNetV2 and DeepLIFT could not be run on ViT.

  • Neuron category distributions vary by architecture. ResNet-18 shows 22.44% ± 0.35 detrimental, 1.48% ± 0.11 neutral, and 76.08% ± 0.39 critical neurons; MobileNet shows 12.29% ± 0.21, 8.35% ± 0.56, and 79.36% ± 0.68; ResNet-50 shows 7.14% ± 0.08, 27.37% ± 0.37, and 65.49% ± 0.38; ViT shows 23.18% ± 0.13, 32.27% ± 0.14, and 44.55% ± 0.27. The authors attribute the higher neutral fractions in ViT and ResNet-50 to over-parameterization.

  • Ablations support the design choices. Neuron evaluation order has negligible impact; the significance level α governs how many neurons are classified as neutral; and a "general inference" variant that compares score distributions across all classes at once flags too many neurons as neutral or detrimental (false negatives), causing overly aggressive pruning and loss of critical information.

  • Segmentation results follow the same pattern. On DDRNet 23S / Cityscapes, C-SWAP reaches 36.9 ± 0.44 SAUCE versus 22.31 ± 0.78 for Neutral, 17.9 ± 0.0 for OMP, 15.0 ± 0.0 for AMP, 10.3 ± 1.2 for Random, 5.24 ± 0.32 for Detrimental, and 72.45 ± 0.57 for Critical. C-SWAP maintains performance up to a 40% pruning ratio, beyond which mIoU drops sharply.

  • Computational cost is measurable but bounded. On a single Nvidia 3090 GPU with 128 samples per class (350 samples total in segmentation), a full C-SWAP run took 1:58 for ResNet-18/CIFAR10 (2,880 neurons explored), 1:02:18 for MobileNetV2/ImageNet (9,128 neurons), 3:24:43 for DDRNet/Cityscapes (3,840 neurons), 7:41:07 for ResNet-50/ImageNet (19,008 neurons), and 34:41:08 for ViT/ImageNet (46,080 neurons).

  • Pruning percentage is a good proxy for compression. The paper reports a near-linear correlation between pruning percentage and both size reduction (number of parameters) and FLOPs reduction, with MobileNet less linear due to its depth-wise convolution layers.

Methodology in Plain English

The authors start from a pre-trained network and a small set of labeled examples (128 per class in classification, 350 images total in segmentation). For each neuron, processed layer by layer from the output back toward the input, they disconnect that neuron's outgoing weights and measure how much the model's score changes.

The score is a task-appropriate performance metric — the probability of the correct class for classification, and mean IoU for segmentation. Averaging the relative change across samples gives a signed number called the global causal effect. A positive value means the neuron hurts performance; a negative value means it helps.

To decide whether a change is real or just noise, they run a paired t-test per class at a 5% significance level, producing a yes/no predicate for each neuron-class pair. These predicates are then combined by voting: a neuron is neutral if no class shows a significant effect, critical if any class shows one and the overall effect is beneficial, and detrimental if any class shows one and the overall effect is harmful.

Pruning then follows directly from this classification. Non-critical neurons are deleted on the spot as the analysis progresses, so the network being tested is always the already-pruned one — this is what makes the procedure adaptive without requiring a global ranking. Critical neurons are kept, and only ranked among themselves at the end in case further pruning is needed. The authors argue this captures most of the benefit of a greedy prune-one-and-rerank loop at a fraction of the cost.

Evaluation uses two tools: pruning curves, which plot validation accuracy against the percentage of parameters removed, and SAUCE, which collapses such a curve into a single number by estimating the area under it. Results are averaged over five random seeds. Baselines include random pruning, one-shot magnitude pruning (OMP), AMP (an adaptation of ACDC), various attribution methods run through Captum, and C-BP, a non-progressive variant of their own method.

Why This Matters

The paper strengthens the case that explainability can be a practical tool for compression rather than only a diagnostic one. It also argues that causal criteria are more robust than correlation- or magnitude-based rankings when a large fraction of a network is removed, and it introduces SAUCE as a way to compare pruning criteria with a single number instead of eyeballing curves.

Real-world applications:

  • Edge deployment of vision models on devices with tight memory and compute budgets, which is the motivation the paper states explicitly.
  • Semantic segmentation for scene understanding, the dense-prediction task the authors demonstrate on Cityscapes with DDRNet.
  • Sustainable or "green" machine learning, cited by the authors as a reason to avoid repeated retraining cycles.
  • Any deployment pipeline where retraining is impractical — for instance when data is unavailable, restricted, or the model is already in production.

Industry relevance: The work comes from Thales (cortAIx Lab, Montreal), and the code is released publicly. Because C-SWAP requires no fine-tuning and needs only a small sample budget, it fits situations where a trained model must be shrunk quickly without access to a full training setup. The computational cost profile matters here too: runs range from under two minutes on ResNet-18/CIFAR10 to over 34 hours for ViT/ImageNet, still far below full retraining and with no post-pruning fine-tuning required.

Future Directions

  • Scaling to very wide layers. The authors state that per-unit causal relevance is tractable for medium-width layers (2048 neurons, roughly 50 minutes) but becomes demanding for very wide ones, and propose group-wise block or channel relevance scores as a mitigation.

  • Extending beyond classification and segmentation. Object detection is named as an open avenue, and the authors note that developing effective scoring functions for object-level tasks would benefit both explainability and compression research.

  • More efficient metrics for dense prediction. Because IoU must be recomputed for every analyzed neuron, DDRNet shows the longest per-neuron time in the study; designing cheaper segmentation metrics is flagged as future work.

  • Open questions implied by the design. How sensitive the critical/neutral/detrimental split is to the choice of α across architectures, and whether the voting strategy can be refined to reduce false positives without losing single-class information, remain areas the paper opens rather than closes. The impact of sample size M is examined for ResNet-18 on CIFAR-10 in an appendix, where the authors report that smaller budgets still produce reliable causal inference at some cost in pruning precision.

Target Audience

This paper is most useful to researchers and engineers working on model compression, efficient inference, and edge deployment, particularly those already familiar with structured pruning and interested in attribution-based or causal criteria. It also suits XAI researchers who want to see explanations used as a decision mechanism rather than as a visualization, and practitioners who need to shrink pre-trained CNNs or vision transformers without access to fine-tuning. Readers unfamiliar with pruning baselines or t-test-based significance testing will need some background, but the method itself is described procedurally enough to follow.

Authors’ abstract

Neural network compression has gained increasing attention in recent years, particularly in computer vision applications, where the need for model reduction is crucial for overcoming deployment constraints. Pruning is a widely used technique that prompts sparsity in model structures, e.g. weights, neurons, and layers, reducing size and inference costs. Structured pruning is especially important as it allows for the removal of entire structures, which further accelerates inference time and reduces memory overhead. However, it can be computationally expensive, requiring iterative retraining and optimization. To overcome this problem, recent methods considered one-shot setting, which applies pruning directly at post-training. Unfortunately, they often lead to a considerable drop in performance. In this paper, we focus on this issue by proposing a novel one-shot pruning framework that relies on explainable deep learning. First, we introduce a causal-aware pruning approach that leverages cause-effect relations between model predictions and structures in a progressive pruning process. It allows us to efficiently reduce the size of the network, ensuring that the removed structures do not deter the performance of the model. Then, through experiments conducted on convolution neural network and vision transformer baselines, pre-trained on classification tasks, we demonstrate that our method consistently achieves substantial reductions in model size, with minimal impact on performance, and without the need for fine-tuning. Overall, our approach outperforms its counterparts, offering the best trade-off. Our code is available on GitHub.

Read the original paper