Research
Scalable Machine Learning Analysis of Parker Solar Probe Solar Wind Data
Overview Research area: Machine learning applied to heliophysics — specifically scalable density estimation and anomaly detection on in-situ solar wind measurements from NASA's Parker Solar Probe (PSP

- arXiv
- 2510.21066
- Published
- 2025-10-24
- Authors
- Daniela Martin, Connor O'Brien, Valmir P Moraes Filho, Jinsu Hong, Jasmine R. Kobayashi, Evangelia Samara, Joseph Gallego
AI summary
Overview
Research area: Machine learning applied to heliophysics — specifically scalable density estimation and anomaly detection on in-situ solar wind measurements from NASA's Parker Solar Probe (PSP).
Technical level: Intermediate. The paper assumes some familiarity with probability distributions (PDFs, CDFs, kurtosis) and with distributed computing concepts, but the machine learning method is described at a conceptual level.
Scope in one sentence: The paper introduces a Dask-plus-Kernel-Density-Matrices pipeline for processing over 150 GB of PSP solar wind data and uses it to derive univariate and bivariate distributions of solar wind speed, proton density, and proton thermal speed as a function of distance from the Sun.
What This Paper Is About
Parker Solar Probe has returned years of measurements of the plasma environment near the Sun, but the resulting dataset (2018–2024, exceeding 150 GB) is too large for conventional analysis tools such as kernel density estimation, which becomes computationally prohibitive at that scale. The authors build a distributed framework that pairs Dask for large-scale statistics with the quantum-inspired Kernel Density Matrices (KDM) method for density estimation, so that the underlying distributions of solar wind parameters — and anomaly thresholds for each — can be estimated directly from the full dataset rather than from binned or averaged subsets. The goal is a tractable, interpretable, and reproducible way to characterize solar wind behavior in the inner heliosphere.
Key Contributions
- Statistical characterization of solar wind properties using distributed processing. The authors compute statistics and distributions from the full PSP/SWEAP dataset spanning 2018 to 2024, organized as a function of heliocentric distance from 0 to 1 AU in increments of 0.1 AU.
- Application of KDM to multi-parameter distributions in large PSP data. KDM is used to produce univariate distributions (e.g., solar wind speed) and bivariate distributions (e.g., solar wind speed versus proton density) at different heliocentric distances, plus anomaly thresholds for each parameter.
- Open-source tools for large in-situ datasets. The code and configuration files are released publicly at https://github.com/spaceml-org/PSP-KDM, along with processed data products, to support reproducible analysis.
Main Findings
- Full-dataset ranges and medians. Boxplots covering 2018 to 2024 give solar wind speeds from 100 km/s to 1980 km/s (likely instrument artifacts at the extremes) with a median of 423 km/s; proton density from 0.00657 to 2,560 cm⁻³ with a median of 734 cm⁻³ (the lower bound possibly a "disappearing solar wind" event); and proton thermal speed (wp_fit) from 4 km/s (0.084 eV, a potential outlier) to 109,900 km/s (63 MeV, almost certainly instrument error) with a median of 95.4 km/s (46 eV).
- Speed increases with distance from the Sun. The cumulative distribution function of solar wind speed, shown from 0.1 to 0.6 AU, indicates that speeds are lower closer to the Sun, consistent with coronal measurements before acceleration.
- Structure changes with distance. Between 0.2 AU and 0.4 AU speeds appear relatively uniform, while at 0.3 AU both fast and slow solar streams emerge.
- Distribution shape varies by distance. The probability density function is right-skewed at all distances; it is leptokurtic at 0.2 AU (relatively homogeneous velocities) and platykurtic at 0.3 AU (broader variations), and at 0.3 AU there is a large high-speed tail above 450 km/s.
- Inverse speed–density relationship. Bivariate distributions from 0.3–0.4 AU and 0.4–0.5 AU both show a clear hyperbolic inverse relationship: as solar wind speed rises, proton density falls. Considerable density variability appears when speed is between 250 and 350 km/s, which the authors suggest may indicate multiple sources or types of solar wind within the slower population.
- Rarefaction with distance. Between the 0.3–0.4 AU and 0.4–0.5 AU bands, the maximum observed solar wind speed falls from approximately 700 km/s to around 600 km/s, and the maximum proton density falls from 200 cm⁻³ to 120 cm⁻³, consistent with rarefaction as the solar wind expands radially.
- KDM hyperparameters. The number of components was set to 400, 800, or 1600, chosen empirically as a balance between flexibility and physical interpretability; the Gaussian kernel width σ was 0.1 and trainable; the learning rate was 10⁻³. Fewer components tend to underfit (smoother densities, slightly higher mean error, better generalization), while too many produce spikier densities that overfit.
- Coverage limits of the analysis. Although analyses covered 0 to 1 AU, the univariate speed distributions are shown only up to 0.6 AU because beyond that point sparse measurements produce spiky, statistically non-robust curves.
Methodology in Plain English
The authors start from raw SWEAP particle measurements, which are distributed as CDF files, and convert them to Zarr format so the data can be processed in parallel across many cores and nodes. Data gaps (fill values) are flagged and removed, and each record is annotated with PSP's radial distance from the Sun. The authors use plasma parameters derived by fitting Gaussians to the measured particle distributions rather than by taking moments, because the fit technique has better temporal coverage.
The analysis runs in two stages. First, Dask computes summary statistics over the whole dataset. Second, the KDM method learns a probability density: it represents the data as a weighted mixture of components, each associated with a Gaussian kernel, and trains the component locations, mixture weights, and kernel width by maximizing the log-likelihood of the observed points using automatic differentiation. This yields explicit univariate and joint distributions, from which anomaly thresholds can be derived. All analyses are performed as a function of heliocentric distance from 0 to 1 AU in 0.1 AU increments. Experiments ran on Google Cloud Platform on a c2-standard-8 VM (8 vCPUs, 32 GB RAM) with 1 TB SSD storage and an Nvidia L4 GPU (24 GB).
Why This Matters
Impact on research. The work shows that a quantum-inspired density estimator can be paired with a distributed data framework to analyze hundreds of gigabytes of in-situ space physics measurements without resorting to binning or averaging that discards structure. It produces explicit density functions (unlike implicit generative models such as GANs or diffusion models) that are easy to interpret physically, and it provides reproducible code and processed data products for the heliophysics community.
Real-world applications.
- Space weather forecasting, where solar wind speed and density measurements feed predictive models.
- Understanding and anticipating geomagnetic storms, which solar wind structures can trigger.
- Studying coronal mass ejections and other extreme solar events that drive hazardous space weather.
- Benchmarking physical models of the inner heliosphere against PSP observations to increase confidence in forecasting frameworks.
Industry relevance. The pipeline pattern — distributed ingestion of large scientific archives plus a trainable, interpretable density model — transfers to any domain with high-volume time series and a need for anomaly thresholds. The work was carried out as part of Heliolab, an initiative of the Frontier Development Lab, a public–private partnership between NASA, Trillium Technologies, and commercial AI partners including Google Cloud and NVIDIA, and it acknowledges Google Cloud computational resources enabled through VMware.
Future Directions
- Extend the statistically robust analysis beyond 0.6 AU. The authors explicitly limit the univariate speed results to 0.6 AU because farther-out measurements are sparse and noisy; recovering reliable distributions at larger heliocentric distances is an open problem.
- Characterize the multi-source slow wind. The density variability observed between 250 and 350 km/s suggests multiple sources or types of solar wind within the slow population; identifying and separating them is a natural next step.
- Validate and apply the anomaly thresholds. The abstract states that anomaly thresholds for each parameter were computed, but no numeric thresholds are reported in the main text, so their practical use for flagging extreme events remains to be demonstrated.
- Broaden to other parameters and instruments. The main text covers solar wind speed and proton density, with other parameters and combinations deferred to supplemental material; extending to the other PSP instruments (FIELDS, WISPR, IS⊙IS) and to other high-volume space physics datasets is presented as a goal of the method's adaptability.
Target Audience
Heliophysics and space physics researchers who work with Parker Solar Probe or other large in-situ datasets; machine learning practitioners interested in scalable, interpretable density estimation and anomaly detection on scientific time series; and space weather researchers who need statistical summaries of solar wind speed and proton density as a function of distance from the Sun.
Authors’ abstract
We present a scalable machine learning framework for analyzing Parker Solar Probe (PSP) solar wind data using distributed processing and the quantum-inspired Kernel Density Matrices (KDM) method. The PSP dataset (2018--2024) exceeds 150 GB, challenging conventional analysis approaches. Our framework leverages Dask for large-scale statistical computations and KDM to estimate univariate and bivariate distributions of key solar wind parameters, including solar wind speed, proton density, and proton thermal speed, as well as anomaly thresholds for each parameter. We reveal characteristic trends in the inner heliosphere, including increasing solar wind speed with distance from the Sun, decreasing proton density, and the inverse relationship between speed and density. Solar wind structures play a critical role in enhancing and mediating extreme space weather phenomena and can trigger geomagnetic storms; our analyses provide quantitative insights into these processes. This approach offers a tractable, interpretable, and distributed methodology for exploring complex physical datasets and facilitates reproducible analysis of large-scale in situ measurements. Processed data products and analysis tools are made publicly available to advance future studies of solar wind dynamics and space weather forecasting. The code and configuration files used in this study are publicly available to support reproducibility.