Research
Physics-guided Neural Network-based Shaft Power Prediction for Vessels
Overview Research area: Machine learning applied to maritime engineering — specifically physics-guided neural networks for predicting vessel shaft power. Technical level: Intermediate. The paper assum

- arXiv
- 2512.20348
- Published
- 2025-12-23
- Authors
- Dogan Altan, Hamza Haruna Mohammed, Glenn Terje Lines, Dusica Marijan, Arnbjørn Maressa
AI summary
Overview
- Research area: Machine learning applied to maritime engineering — specifically physics-guided neural networks for predicting vessel shaft power.
- Technical level: Intermediate. The paper assumes familiarity with neural network training, loss functions, and standard regression metrics, but the physical formulas are explained explicitly.
- Scope: The paper introduces a hybrid neural network that embeds empirical calm-water, wind, and wave resistance formulas into its training loss, and evaluates it against a pure empirical-formula method and a plain neural network on data from four cargo vessels.
What This Paper Is About
Accurately predicting a ship's shaft power matters because shaft power is closely tied to fuel consumption, cost, and emissions. Traditional empirical formulas model resistance well in steady conditions but struggle with dynamic factors such as sea state and hull fouling, while pure machine learning models can produce physically implausible outputs. The authors' goal is to combine both: a neural network whose training is guided by domain physics so that predictions remain consistent with marine engineering principles while still adapting to real-world data.
Key Contributions
- A physics-guided neural network (PGNN) model for vessel shaft power prediction in which physical formulas are incorporated directly into the model's loss function, injecting domain-specific insight into the learning process.
- Inclusion of three resistance components in the training signal — calm water, wave, and wind resistance — rather than calm water alone, giving a more complete representation of environmental factors.
- A two-stage prediction pipeline in which shaft rotational speed (RPM) is first predicted with a polynomial model and then supplied as an input feature to the shaft power network, avoiding reliance on ground-truth RPM so the comparison with the empirical-formula baseline is fair.
- An evaluation on real-world data from four similar-sized cargo vessels against two baselines, plus a dedicated analysis of performance under severe sea conditions (high waves) on an Atlantic crossing.
Main Findings
- PGNN beats both baselines across all four vessels on error metrics: The physics-guided network achieved lower mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE) than both the empirical formula method (EF) and the base neural network (NN) for Vessels A, B, C, and D.
- Representative Vessel A numbers: PGNN reached MAE 761.12 ± 109.69, RMSE 972.52 ± 130.46, R² 0.498 ± 0.12, and MAPE 14.50 ± 2.33, compared with NN (MAE 837.14 ± 134.06, MAPE 15.80 ± 2.44) and EF (MAE 1450.10 ± 17.39, MAPE 26.28).
- Empirical formulas fit poorly overall: EF produced R² scores that were either close to zero (Vessel C, 0.011) or negative (Vessel A −0.483, Vessel B −2.307, Vessel D −0.195), indicating poor fit to the data. PGNN's R² was positive for every vessel.
- One metric-level exception: For Vessel D, EF yielded a better MAE (799.83 ± 7.63) than NN (819.80 ± 113.03), though NN was better on the other metrics — the paper notes this as an observation rather than a general pattern.
- Over-time predictions: The model's predictions were mostly consistent with actual shaft power for Vessel A, though offset patterns appeared from time to time after July 2023. Predictions for Vessel B largely overlapped the actual values. Vessel C showed some shifts, for example around June 25. Vessel D deviated more, consistent with its higher error scores, but predictions became more accurate after the propeller polish event.
- Strong performance in rough seas: On an Atlantic crossing, PGNN achieved the best scores on all reported metrics for both Vessel A and Vessel E. For Vessel A the MAE dropped to 551.70 ± 63.01 (from 759.00 ± 133.74 for NN and 1257.02 ± 1.27 for EF), and MAPE to 7.00 ± 0.89. For both vessels PGNN gave almost 2% improvement in MAPE compared to NN on average.
- RPM prediction accuracy: The polynomial model used to predict RPM achieved an average MAPE of 3.95%. Removing the predicted RPM as a feature decreased MAPE by approximately 1.6%.
- Fouling is handled as a feature, not as physics: The physics-guided component of the method lacks a representation of fouling; fouling effects enter only as input features such as days elapsed after propeller polish and dry docking.
Methodology in Plain English
The authors work with sensor data sampled every 15 minutes from four cargo vessels, all between 204 and 208 meters long, plus environmental data from Copernicus. The input features are speed through water, draught, sea depth, sea temperature, wave height, swell height, wave direction, swell direction, wind direction, wind speed, days since propeller polish, and days since dry docking. All direction features are relative to the vessel.
They construct three resistance formulas adapted from International Towing Tank Conference guidelines: a calm-water resistance formula with learnable coefficients, a wind resistance formula with learnable cylinder, headwind, and sidewind factors, and a wave resistance formula using a fixed constant of 2.5 for the wave exponent plus a learned geometric factor. Summing the three resistances and multiplying by speed through water yields a physically estimated power.
For the neural network, features are split into groups (Copernicus-derived, sensor, and vessel-event features) and each group passes through its own dense layers before being concatenated and passed through further dense layers with dropout. The key twist is in training: the loss is the usual prediction error plus a second term measuring the gap between the network's prediction and the physically estimated power, scaled by a weighting factor lambda. The authors trained models across lambda values from 0.05 to 1 and selected the best on the test set, ending up with 0.1, 0.1, 0.05, and 0.4 for Vessels A, B, C, and D.
Because the empirical baseline does not use RPM, the authors predict RPM with a multiplicative polynomial model (four features — speed through water, draught, wind speed, and swell height — each at order 3) rather than using ground-truth RPM. Data were preprocessed by dropping instances below 5 knots or below 500 kW and removing missing values; training used the Adam optimizer, mean absolute error loss, batch size 16, 20% of the training set for validation, and early stopping with patience 10. Each experiment was repeated 10 times, with mean and standard deviation reported.
Why This Matters
For research, this paper is a concrete demonstration that embedding physics into a loss function — rather than using physics only for preprocessing or post-hoc correction — improves generalization for an industrial regression problem. It also shows the limits of empirical formulas when data drift from fouling is present in training and test sets, since those formulas produced negative R² on three of four vessels.
Real-world applications include:
- Voyage planning and speed optimization, where accurate shaft power forecasts feed into fuel and emissions estimates.
- Performance monitoring of hull and propeller condition, since predictions diverge from actuals as fouling accumulates.
- Regulatory reporting and decarbonization planning, aligning with the International Maritime Organization's 2023 greenhouse gas reduction strategy.
- Dry docking and maintenance scheduling, where fouling-driven efficiency losses exceeding 10% can justify maintenance timing decisions.
For industry, the relevance is direct: maritime transport moves more than 80% of all goods worldwide by volume, so even modest percentage improvements in fuel prediction translate into large cost and emissions reductions across fleets.
Future Directions
- Integrating fouling conditions as part of the physical loss, rather than only as input features, to better learn fouling's effect on shaft power.
- Conducting an ablation analysis of the individual resistance components used in the physical loss to quantify each one's contribution.
- Testing the approach on data obtained from different types of vessels beyond the similar-sized cargo vessels studied here.
- Addressing the observed tendency of the physics-guided network toward significant prediction errors when shaft power increases or decreases instantaneously, and the general limited extrapolation of neural networks to conditions unseen during training.
Target Audience
Machine learning researchers working on physics-informed and physics-guided models, maritime and naval engineering researchers studying ship resistance and propulsion, and data scientists or fleet performance analysts at shipping companies who need reliable shaft power and fuel consumption predictions. Industrial readers should note that the underlying data and code are confidential and belong to the industrial partner, so the paper cannot be reproduced directly from its contents.
Authors’ abstract
Optimizing maritime operations, particularly fuel consumption for vessels, is crucial, considering its significant share in global trade. As fuel consumption is closely related to the shaft power of a vessel, predicting shaft power accurately is a crucial problem that requires careful consideration to minimize costs and emissions. Traditional approaches, which incorporate empirical formulas, often struggle to model dynamic conditions, such as sea conditions or fouling on vessels. In this paper, we present a hybrid, physics-guided neural network-based approach that utilizes empirical formulas within the network to combine the advantages of both neural networks and traditional techniques. We evaluate the presented method using data obtained from four similar-sized cargo vessels and compare the results with those of a baseline neural network and a traditional approach that employs empirical formulas. The experimental results demonstrate that the physics-guided neural network approach achieves lower mean absolute error, root mean square error, and mean absolute percentage error for all tested vessels compared to both the empirical formula-based method and the base neural network.