Research
Explainable AI for microseismic event detection
Overview Research area: Applied machine learning and interpretable AI in geophysics, specifically microseismic event detection and seismic phase picking. Technical level: Intermediate. The reader bene

- arXiv
- 2510.17458
- Published
- 2025-10-20
- Authors
- Ayrat Abdullin, Denis Anikiev, Umair Bin Waheed
AI summary
Overview
Research area: Applied machine learning and interpretable AI in geophysics, specifically microseismic event detection and seismic phase picking. Technical level: Intermediate. The reader benefits from familiarity with deep learning (convolutional networks, gradients, probability outputs) and basic seismology (P-waves, S-waves, three-component recordings), though the paper explains both. Scope: The paper applies Grad-CAM and SHAP to a PhaseNet-based microseismic detector to explain its decisions, then converts those explanations into a SHAP-gated inference rule that improves detection performance.
What This Paper Is About
Deep neural networks such as PhaseNet detect microseismic events with high accuracy but operate as "black boxes," so seismologists cannot easily judge how much to trust an automated detection. The authors ask two questions: which parts of a waveform and which sensor components drive PhaseNet's decisions, and whether that explanation can be used to make the detector itself more reliable. Their goal is to interpret the model against established geophysical principles and then use those explanations to reduce classification errors.
Key Contributions
- Explanation of PhaseNet's detection behavior. The authors apply Gradient-weighted Class Activation Mapping (Grad-CAM) and Shapley Additive Explanations (SHAP) to PhaseNet adapted for binary microseismic event detection (signal vs. noise), revealing which waveform segments and which components the network relies on.
- A tractable SHAP formulation for long waveform windows. Because computing Shapley values for every sample of a seismic waveform is intractable, they reduce the problem to component-level attributions: full enumeration of all 2^3 = 8 coalitions of the East, North, and Vertical (E, N, Z) channels with either the true signal or a zero baseline, giving exact Shapley contributions without approximation.
- A SHAP-gated inference scheme. They define an explanation-based evidence score as the mean of six absolute SHAP values (E, N, Z for both P- and S-wave outputs) and accept a detection only when this score exceeds a threshold calibrated on the training set. They describe this as the first instance in seismic event detection where explanations inform the model's output in a closed-loop (post-hoc decision fusion) fashion.
- Implemented code release. The implementation and scripts are stated to be publicly available at https://github.com/ayratabd/xAI_PhaseNet.
Main Findings
- Grad-CAM attention aligns with P- and S-wave arrivals. For a high-SNR event, activations are sharply concentrated around the manually picked P arrival, with a weaker secondary highlight at the S arrival, and importance scores extend across several tens of samples around the onset. Outside these arrival times Grad-CAM values are much smaller. For a low-SNR event, attributions still correspond to the P- and S-wave neighborhoods but are broader and less sharply defined. For a noise-only window, activation is weak and scattered with no clustering near any onset.
- Vertical component drives P-phase picks; horizontals drive S-phase picks. In SHAP summary statistics (Table 1), the Z component has the largest effect for P-class detections on signal windows with a mean absolute SHAP value of 0.30 and dominance in 49.7% of cases, while E and N are more important for S-class detections with mean absolute SHAP values of 0.36 and 0.33 and dominance frequencies of 50.5% and 45.6%. This matches the geophysical expectation that S-wave energy is mostly recorded on horizontals.
- SHAP magnitude clearly separates signal from noise. Signal windows show mean absolute SHAP values of about 0.30-0.33 for the P-class and 0.19-0.36 for the S-class, whereas noise windows remain much lower at about 0.06-0.10 and 0.02 respectively. Noise histograms concentrate near zero and overlap strongly across components, showing no stable phase-consistent attribution pattern.
- Component-level noise attributions are an order of magnitude weaker. For noise windows, mean SHAP values are 0.06 (P-class) and 0.02 (S-class). The Z component shows slightly higher noise attributions for P-class (0.10) but these remain an order of magnitude weaker than for true events.
- SHAP-gated inference improves detection on the test set. On a test set of 9,000 waveforms, the SHAP-gated model achieved an F1-score of 0.98 (precision 0.99, recall 0.97), outperforming the baseline PhaseNet (F1-score 0.97). The paper reports that the scheme improved F1 from 0.97 to 0.98 and increased recall from 0.96 to 0.97, while also showing greater robustness under progressively stronger noise contamination.
- Baseline thresholding was problematic. The pre-trained PhaseNet achieved roughly 97% classification accuracy on the held-out test set, but a simple 0.5 probability cutoff was not satisfying, and tuning the threshold involved trading off false negatives against false positives.
- Pairwise interaction analysis. The authors extended the framework to pairwise Shapley Interaction Indices, measuring whether two components act synergistically (positive) or redundantly (negative), computed exactly from the same 8 coalition values. Specific numeric interaction results are not detailed in the available content.
Methodology in Plain English
The study uses triggered three-component waveforms recorded during hydraulic fracturing operations in British Columbia, Canada, from nine surface seismic sensors spread over about 100 km². The original 250 Hz recordings were downsampled to 100 Hz (sufficient because the dominant signal frequency is below 50 Hz). The events are low-magnitude induced events (0.5 < M < 2.5) at an average depth of 2.1 km (range 1.7 to 2.4 km), with manually picked P- and S-wave arrival times. Every waveform window is 30 seconds long and labeled binary: signal if it contains a microseismic arrival, noise otherwise. Noise windows come from continuous recordings during intervals without detected seismicity before operations began, and include field-specific non-seismic background and transient noise. The curated balanced dataset has approximately 10,000 windows; for each run, 100 windows were randomly selected as a balanced training subset for threshold tuning, and evaluation was on a separate test set of 9,000 windows drawn from the remaining 9,900. For the harmonic-noise robustness experiments, injected harmonic noise came from a separate field dataset acquired in Spain with five surface stations and strong pump-induced harmonic noise.
The detection model is PhaseNet, a U-Net style fully convolutional encoder-decoder originally designed for picking P and S arrival times. Its input is 3 × 3001 (three components, 30 seconds at 100 Hz); it has four down-sampling stages using 1-D convolutions with kernel size 7 and stride 4, four up-sampling stages using deconvolutions, skip connections between matching stages, and a final softmax producing three probability sequences (P-pick, S-pick, noise). For binary detection, the authors take the maximum predicted likelihood of a P or S arrival in the window as the event probability.
To explain the model, they first adapted Grad-CAM to one-dimensional time series: gradients of the event output with respect to the final convolutional layer's feature maps are averaged over time positions to weight each filter, a weighted combination of feature maps is passed through a ReLU to keep positive influences, and the resulting coarse activation map is linearly interpolated back to the original waveform length to form an importance score over time. The authors note this emphasizes high-level semantic features at the cost of spatial resolution and that standard Grad-CAM can miss negatively contributing features, which they accept because they focus on positive evidence supporting a detection.
For SHAP, they avoid per-sample Shapley computation by masking whole components. For each window they evaluate all 2^3 = 8 coalitions of E, N, Z (masked channels replaced by zeros), where the score of a coalition is the maximum softmax probability of the target class over the window. With the Shapley weights w(0) = 1/3, w(1) = 1/6, w(2) = 1/3, they compute exact channel attributions for the P and S classes. Per-channel importance over a batch is reported as the mean absolute Shapley value, and they also record how often each component is dominant.
For SHAP-gated inference, they define two scalars per window: the PhaseNet event-probability score (the maximum P-or-S probability in the window) and the SHAP evidence statistic S6, the mean of the six absolute SHAP values (E, N, Z for P and S). Thresholds for the probability rule and the SHAP rule were optimized on the training set by sweeping values to maximize F1-score, then fixed and applied to the test set. The thresholds are dimensionless but tied to the data normalization and model output scaling, so they would need recalibration in another setting.
Why This Matters
The paper argues that XAI can do more than explain a black-box model: it can be folded into the decision rule and directly improve reliability. By showing that PhaseNet's attributions agree with physical wave propagation (Z dominance for P, horizontal dominance for S) and that noise windows lack coherent explanatory evidence, the work offers a way to build justified trust in automated seismic detectors in high-stakes geoscience settings.
Real-world applications:
- Induced seismicity monitoring for CO2 storage or mining, where reliability and transparency are described as crucial.
- Hydraulic fracturing monitoring, the source setting of the dataset used here, where automated event screening supports operational decisions.
- Analyst-assisted event cataloging, where an explanation-based evidence score can help analysts decide which automated triggers to accept or reject, reducing missed detections and improving confidence in automated triggers.
- Detector deployment under noisy conditions, since the SHAP-gated scheme showed greater robustness under progressively stronger noise contamination.
Industry relevance: the authors frame this as supporting safer deployment of AI-assisted decision-making systems in geophysical monitoring, and note this approach reflects a developing trend in which XAI is used not only for interpretation but also for enhancing performance, citing related work in seismic denoising, ground-motion parameter prediction, and building damage detection.
Future Directions
- Extending the explainability-guided strategy to other architectures. The Discussion section considers how the same approach could be extended to emerging models, including Transformer-based detectors.
- Broader deployment of geophysical AI models. The authors consider how XAI may help enable wider deployment of geophysical AI models by addressing reliability concerns raised by earlier researchers.
- Threshold calibration and transferability. Because the thresholds are dimensionless but tied to the data normalization and model output scaling, they would need recalibration in a different setting; establishing how to recalibrate or generalize them is an open practical question.
- Richer attribution methods. The authors note that standard Grad-CAM misses negatively contributing features and mention guided backpropagation or LRP as options to capture inhibitory factors, and cite alternative strategies such as optimal layer selection and fusion of heatmaps from multiple layers, all of which remain outside the scope of this study.
Target Audience
Seismologists and microseismic monitoring practitioners who rely on automated detectors; machine learning researchers working on interpretability for time-series and scientific data; and geophysicists evaluating whether AI-based phase picking can be trusted in operational, safety-relevant settings. Readers with a background in deep learning and seismic signal processing will get the most from the technical detail, while the framing and decision-fusion idea are accessible to a broader applied audience.
Authors’ abstract
Deep neural networks like PhaseNet show high accuracy in detecting microseismic events, but their black-box nature is a concern in critical applications. We apply Explainable Artificial Intelligence (XAI) techniques, such as Gradient-weighted Class Activation Mapping (Grad-CAM) and Shapley Additive Explanations (SHAP), to interpret the PhaseNet model's decisions and improve its reliability. Grad-CAM highlights that the network's attention aligns with P- and S-wave arrivals. SHAP values quantify feature contributions, confirming that vertical-component amplitudes drive P-phase picks while horizontal components dominate S-phase picks, consistent with geophysical principles. Leveraging these insights, we introduce a SHAP-gated inference scheme that combines the model's output with an explanation-based metric to reduce errors. On a test set of 9,000 waveforms, the SHAP-gated model achieved an F1-score of 0.98 (precision 0.99, recall 0.97), outperforming the baseline PhaseNet (F1-score 0.97) and demonstrating enhanced robustness to noise. These results show that XAI can not only interpret deep learning models but also directly enhance their performance, providing a template for building trust in automated seismic detectors. The implementation and scripts used in this study will be publicly available at https://github.com/ayratabd/xAI_PhaseNet.