Research
Domain-Decomposed Graph Neural Network Surrogate Modeling for Ice Sheets
Domain-Decomposed Graph Neural Network Surrogate Modeling for Ice Sheets Overview Research area: Scientific machine learning — specifically physics-inspired graph neural network (GNN) surrogate modeli

- arXiv
- 2512.01888
- Published
- 2025-12-01
- Authors
- Adrienne M. Propp, Mauro Perego, Eric C. Cyr, Anthony Gruber, Amanda A. Howard, Alexander Heinlein, Panos Stinis, Daniel M. Tartakovsky
AI summary
Domain-Decomposed Graph Neural Network Surrogate Modeling for Ice SheetsOverview
- Research area: Scientific machine learning — specifically physics-inspired graph neural network (GNN) surrogate modeling for partial differential equation (PDE) simulations, applied to ice sheet dynamics.
- Technical level: Advanced. The paper assumes familiarity with PDEs, finite element/finite volume discretizations, Hamiltonian systems, graph neural networks, Bayesian inference, and uncertainty quantification.
- Scope: The authors develop a GNN surrogate that predicts ice sheet velocity on an unstructured mesh of Greenland's Humboldt Glacier, and introduce transfer learning and domain decomposition strategies to make training faster and more accurate at large scale.
What This Paper Is About
High-fidelity ice sheet simulators are accurate but slow, and tasks such as uncertainty quantification (UQ) require hundreds or thousands of model evaluations, which makes them computationally intractable with traditional solvers. The authors build a surrogate model that learns to predict the full-field ice velocity field directly from inputs such as basal friction, bed topography, and ice thickness, without remeshing or interpolation. To make this feasible at the scale of a real glacier — a mesh with tens of thousands of nodes — they decompose the domain into subdomains, train local GNN surrogates in parallel, and use transfer learning to fine-tune across subdomains.
Key Contributions
- A physics-inspired GNN surrogate that accurately models ice sheet velocity on a large, unstructured mesh, using a bracket-based (energy-conserving Hamiltonian) message-passing architecture that mitigates oversmoothing.
- A transfer learning strategy that accelerates GNN training by reusing learned attention mechanisms across parts of the domain or related graphs.
- A domain decomposition (DD) framework for GNNs, partitioning the mesh into subdomains, training local surrogates in parallel, and aggregating their predictions — the authors state this has not previously been extended to GNNs.
- A combined DD + transfer learning pipeline, which the authors present as a scalable pathway for training GNN surrogates on massive PDE-governed systems, with a foundation for future UQ studies.
Main Findings
- Full-field velocity prediction works: The abstract reports that the approach "accurately predicts full-field velocities on high-resolution meshes." Specific error metrics are not reported in the available paper content.
- Training time reduction: The domain-decomposed approach "substantially reduces training time relative to a single global surrogate model." The specific runtime figures and speedups are not reported in the available content (the paper references runtime details in Appendix A, which was not included).
- Data-rich ice sheet testbed: Each MALI simulation generates 93 annual snapshots of ice sheet state variables; the authors use each snapshot as an independent training sample, randomly selecting 20 simulations for validation and 20 for testing.
- Graph size: The dual Delaunay triangulation used as the GNN graph has 18,544 nodes and 54,962 edges.
- Higher-fidelity data than prior work: The Humboldt Glacier mesh resolution is roughly eight times finer than in the prior work referenced as [16]; the model also adds calving, basal melting, and temperature evolution, and uses Budd's nonlinear sliding law rather than a linearized form.
- Architecture choices mattered less than features: Experiments indicated that variations in brackets, ODE integrators, and encoding/decoding strategies were far less impactful than feature design and hyperparameter tuning (e.g., the learning rate scheduler), though they did affect training efficiency. The energy-conserving Hamiltonian bracket performed best across experiments.
- No physics-based regularization used: The loss contains no explicit mass-conservation penalties; the authors state the architecture plus the rich feature set was sufficient for stable, accurate predictions.
Methodology in Plain English
The authors start with a physics simulator — MPAS-Albany Land Ice (MALI) — run on Greenland's Humboldt Glacier, spanning nearly 2 million square kilometers. MALI solves ice thickness and temperature on a Voronoi mesh and velocity on its dual Delaunay triangulation using low-order Lagrangian finite elements. The authors adopt that dual triangulation as the graph their GNN operates on, so the machine learning model fits naturally onto the mesh with no remeshing or interpolation.
To generate training data, they sample basal friction fields from a posterior distribution. Basal friction cannot be measured directly, so they use a Laplace approximation around the maximum a posteriori (MAP) estimate to get a tractable Gaussian posterior, draw samples as p = p_MAP + Lω with ω drawn from a standard normal, and set μ = exp(p) to keep friction positive. They simulate forward from year 2007 to year 2100.
Each graph node carries a seven-dimensional feature vector: ice thickness, bed topography (elevation), basal friction, and two Boolean indicators for grounded versus floating ice. The model outputs the x- and y-components of velocity. Features are z-score normalized using global statistics computed once and reused across all experiments, while edge-wise distances are min-max normalized to [0,1].
The core network is a bracket-based GNN in which message passing is reformulated as a Hamiltonian dynamical system rather than diffusive averaging. Information passes through three phases: an encoding phase that lifts raw node-edge features into a latent space; a message-passing phase where latent features evolve in a pseudo-time T according to an autonomous graph neural ODE governed by a skew-adjoint operator, which guarantees the energy rate is exactly zero and prevents homogenization of features; and a decoding phase that maps the final latent state back to physical predictions. Graph attention is implemented through learnable inner product matrices — the edge weights are exponentials of symmetrized query-key dot products, and the node inner product is the sum of incident edge weights — yielding a graph-attention-like Laplacian that encodes both topology and learned metric information.
Training minimizes mean squared error between predicted and simulated velocities over randomly selected time steps and independent basal friction fields, using mini-batch gradient descent on NVIDIA A100 40GB GPUs. The paper then layers on two strategies: domain decomposition (partition the mesh, train local GNNs in parallel, aggregate), and transfer learning (reuse learned weights across subdomains to accelerate training and improve accuracy when data are limited).
Why This Matters
Impact on research: The work shows that GNN surrogates can be scaled to large, unstructured, physics-heavy domains by combining classical domain decomposition ideas with machine learning. Because GNN weights can be transferred across graphs of consistent feature dimension, this offers a route to surrogate modeling that generalizes across mesh resolution, input parameters, and domain geometry — something architectures like standard DeepONets cannot do without retraining. The paper explicitly frames this as groundwork for UQ studies that were previously computationally out of reach.
Real-world applications (as motivated in the paper):
- Coastal infrastructure planning — reliable predictions of ice movement and mass loss underpin decisions about coastal defenses.
- Resource planning — long-horizon projections of glacial change inform planning in affected regions.
- Climate adaptation policy — quantifying risk and exploring plausible scenarios requires the hundreds-to-thousands of evaluations that fast surrogates make possible.
- Uncertainty quantification over unobservable parameters — the basal friction coefficient μ cannot be measured directly, so efficient surrogate modeling is central to characterizing its uncertainty and propagating it into projections.
Industry relevance: Climate-risk analytics, reinsurance and infrastructure risk assessment, and any sector requiring ensembles of geophysical simulations under uncertainty all benefit from surrogate models that reduce evaluation cost by orders of magnitude without abandoning the underlying mesh. The domain decomposition approach also maps naturally onto distributed computing infrastructure, which is directly relevant to large-scale scientific computing workflows.
Future Directions
- Uncertainty quantification studies of ice sheet dynamics — the authors describe their framework as laying the groundwork for UQ objectives, and the paper's concluding section (Section 7) discusses implications and future directions, though the content available here is truncated.
- Application beyond ice sheets — the authors state the framework has "far-reaching implications beyond the context of ice sheet dynamics" and broad potential for application to other PDE-governed systems.
- Generalization across evolving geometry — since ice sheet domains and meshes may evolve, extending transferability of trained GNNs to changing meshes and domain shapes is an open direction. The paper notes the GNN formulation requires no remeshing or geometric preprocessing if the domain or mesh changes.
- Scaling the DD plus transfer learning combination — the abstract frames graph-based domain decomposition combined with transfer learning as "a scalable and reliable pathway for training GNN surrogates on massive PDE-governed systems," leaving further scaling and refinement as an open problem.
Target Audience
This paper is most valuable to researchers and practitioners working at the intersection of scientific machine learning and geophysical modeling: developers of neural surrogate models for PDEs, computational glaciologists and climate scientists needing fast forward models for UQ, and numerical analysts interested in how classical domain decomposition translates to graph-based learning architectures. Applied mathematicians and ML engineers working on graph neural networks for physical simulation will also find the bracket-based attention formulation relevant. A strong background in PDE discretization and deep learning is needed to follow the technical sections in full.
Note on precision: the supplied paper content is truncated partway through Section 4, so quantitative results from the experiments (error metrics, training-time speedups, transfer-learning gains) are not available and are not reported here. All figures cited above — 2 million square kilometers, 18,544 nodes, 54,962 edges, 93 snapshots, 20 validation and 20 test simulations, 8× finer mesh, year 2007 to 2100, seven-dimensional node features, γ = 8.976 km², δ = 8.865 × 10⁻³, ξ = 0.1987 km⁻¹, 90 km correlation length, q = 1/3 — come directly from the provided text.
Authors’ abstract
Accurate yet efficient surrogate models are essential for large-scale simulations of partial differential equations (PDEs), particularly for uncertainty quantification (UQ) tasks that demand hundreds or thousands of evaluations. We develop a physics-inspired graph neural network (GNN) surrogate that operates directly on unstructured meshes and leverages the flexibility of graph attention. To improve both training efficiency and generalization properties of the model, we introduce a domain decomposition (DD) strategy that partitions the mesh into subdomains, trains local GNN surrogates in parallel, and aggregates their predictions. We then employ transfer learning to fine-tune models across subdomains, accelerating training and improving accuracy in data-limited settings. Applied to ice sheet simulations, our approach accurately predicts full-field velocities on high-resolution meshes, substantially reduces training time relative to training a single global surrogate model, and provides a ripe foundation for UQ objectives. Our results demonstrate that graph-based DD, combined with transfer learning, provides a scalable and reliable pathway for training GNN surrogates on massive PDE-governed systems, with broad potential for application beyond ice sheet dynamics.