Research
Seabed-Net: A multi-task network for joint bathymetry estimation and seabed classification from remote sensing imagery in shallow waters
Overview Research area: Computer vision and remote sensing — multi-task deep learning applied to shallow-water seabed mapping, specifically joint spectrally-derived bathymetry (SDB) and pixel-based se

- arXiv
- 2510.19329
- Published
- 2025-10-22
- Authors
- Panagiotis Agrafiotis, Begüm Demir
AI summary
Overview
Research area: Computer vision and remote sensing — multi-task deep learning applied to shallow-water seabed mapping, specifically joint spectrally-derived bathymetry (SDB) and pixel-based seabed classification from satellite and aerial imagery.
Technical level: Advanced. The paper assumes familiarity with encoder-decoder segmentation networks, attention mechanisms (spatial and channel), Swin Transformers, cross-attention, and uncertainty-based multi-task loss weighting.
Scope: The paper proposes and evaluates Seabed-Net, a dual-branch multi-task network that predicts continuous water depth and seabed class labels simultaneously from Sentinel-2, SPOT 6, and airborne RGB imagery, benchmarked on the MagicBathyNet dataset.
What This Paper Is About
Depth estimation and seabed classification from remote sensing imagery are normally treated as separate problems, which discards the fact that each task carries information the other needs — seabed type explains brightness variations that depth models otherwise misread as depth, and depth structure marks the boundaries where seabed classes change. The authors build a single network that performs both tasks at once, sharing features between the two branches, and test whether this joint training beats models that do only one task. The goal is a mapping tool that stays accurate across very different sensors and coastal environments rather than degrading when image resolution gets coarser.
Key Contributions
-
Seabed-Net architecture. The authors propose what they describe as the first dual-branch multi-task architecture for jointly performing bathymetry retrieval and pixel-level seabed classification, where depth cues refine class boundaries and habitat information improves depth estimation.
-
Cross-task fusion design. Two complementary fusion mechanisms are introduced at every encoder scale: an Attention Feature Fusion (AFF) module combining spatial and channel attention to capture local cross-task dependencies, and a ViT-based fusion module using a lightweight Swin Transformer block with window-based multi-head self-attention and cross-attention to capture global context. AFF-derived features feed the bathymetry decoder as skip connections; ViT-derived features feed the classification decoder.
-
Uncertainty-based joint optimization. Instead of hand-tuned fixed loss weights, the framework learns task-specific uncertainty parameters (per Kendall et al., 2018) to dynamically balance the bathymetry RMSE loss against the classification cross-entropy loss.
-
Multi-sensor evaluation and ablation. The method is evaluated across three sensors of different spatial resolutions at two coastal sites, compared against empirical, traditional machine learning, single-task, and multi-task baselines, with additional ablation and cross-modal analyses. Code and pretrained weights are released at https://github.com/pagraf/Seabed-Net.
Main Findings
-
Gains over empirical models: Seabed-Net outperformed Multiband Linear and Log Band Ratio across all datasets and locations in RMSE, MAE, and standard deviation, with RMSE reductions typically ranging from 40% to 60% for aerial data, 60% to 70% for SPOT 6, and 65% to 75% for Sentinel-2. The abstract frames the overall figure as up to 75% lower RMSE against traditional empirical models and traditional machine learning regression methods.
-
Gains over machine learning regressors: Against Random Forest Regression and Support Vector Regression, Seabed-Net reduced RMSE by approximately 24% to 42% and 32% to 49% respectively in Agia Napa, and maintained 14% to 39% lower errors in the more uniform Puck Lagoon. In Sentinel-2 imagery over Puck Lagoon, Seabed-Net reached nearly 0.43 m RMSE versus 0.642 m for Random Forest and 0.637 m for SVR.
-
Gains over the single-task bathymetry model: Compared with UNet-bathy, Seabed-Net reduced RMSE by approximately 10-22% (aerial), 9-19% (SPOT 6), and 24-38% (Sentinel-2); MAE by 12-26% (aerial), 8-39% (SPOT 6), and 27-35% (Sentinel-2); and standard deviation by 10-23% (aerial), 6-44% (SPOT 6), and 25-51% (Sentinel-2).
-
Reported comparison with state-of-the-art baselines: The abstract states that Seabed-Net reduces bathymetric RMSE by 10-30% relative to state-of-the-art single-task and multi-task baselines and improves seabed classification accuracy by up to 8%. The detailed per-model tables for classification and for the multi-task baselines (PAD-Net, MTI-Net, MTL, JSH-Net, TaskPrompter) are not included in the provided text.
-
Site difficulty matters for empirical methods: At Agia Napa, empirical models produced errors up to 2-3 times higher than at the relatively uniform Puck Lagoon, while Seabed-Net kept consistently low error at both sites.
-
Resolution sensitivity: The paper reports that conventional single-task models degrade significantly under lower-resolution satellite data, and positions the multi-task framework as more resilient to resolution degradation, with the largest reported improvements appearing in Sentinel-2 imagery.
-
Qualitative improvements: Analyses report enhanced spatial consistency, sharper habitat boundaries, and corrected depth biases in low-contrast regions.
-
Prediction stability: Seabed-Net showed markedly lower standard deviation values than the regression baselines, which the authors interpret as greater prediction stability.
Methodology in Plain English
The input is an RGB image. Two parallel convolutional encoders process it, one for depth and one for class labels, each producing feature maps at four scales. The bathymetry encoder deliberately omits batch normalization so fine radiometric detail is preserved; the classification encoder keeps batch normalization to help generalization. At each of the four scales, the two branches' feature maps are fused twice: once through an attention module that concatenates them and applies a spatial attention map (a 7x7 convolution followed by a sigmoid) plus a channel attention map from a squeeze-and-excitation block, and once through a transformer module that projects the concatenated features, processes them with a Swin Transformer block using window-based self-attention and cross-attention, then reprojects them.
Decoding is asymmetric. The attention-fused features are added to the bathymetry encoder features and routed to the depth decoder, while the transformer-fused features are added to the classification encoder features and routed to the classification decoder. The depth decoder omits normalization to protect continuous-value precision; the classification decoder applies batch normalization. Training combines a root mean squared error loss for depth and a cross-entropy loss for classes, weighted by learnable uncertainty parameters rather than fixed weights.
Experiments use the MagicBathyNet dataset at Agia Napa in Cyprus and Puck Lagoon in Poland, with co-registered Sentinel-2 Level2A, SPOT 6, and aerial orthoimagery, only RGB bands, and no sun glint removal, using the dataset's predefined 80%-20% train/test splits. Training ran in PyTorch on one NVIDIA A100 80GB GPU with Adam, a cosine annealing scheduler over 10 epochs of 10,000 iterations each (100,000 gradient updates), initial learning rates of 10^-4 for SPOT 6 and Sentinel-2 and 10^-5 for aerial data, 256x256 resized crops for the satellite sensors, and random rotations plus vertical and horizontal flips as augmentation. Evaluation uses RMSE, MAE, standard deviation, overall accuracy, and mean Intersection over Union.
Why This Matters
Impact on research. The paper argues that single-task depth and classification models are brittle, resolution-dependent, and prone to spatially incoherent or over-smoothed output, and that joint modeling with uncertainty-based balancing is a viable remedy. It also extends multi-task learning techniques that are established in terrestrial remote sensing into the shallow-water domain, where the paper states such use remains largely unexplored.
Real-world applications (as motivated by the paper):
- Coastal habitat monitoring and mapping of substrates and benthic habitats such as seagrass, macroalgae, eelgrass, sand, and rock.
- Navigation safety and shallow-water surveying, where echo-sounders suffer from wave-induced interference, reefs, and multi-path propagation errors and airborne LiDAR remains costly.
- Monitoring of submerged cultural heritage and responses to maritime tragedies and natural disasters.
- Offshore energy and marine resource development, plus tracking of habitat destruction and marine pollution under climatological and anthropogenic pressure.
Industry relevance. The approach is designed to work on RGB-only imagery, which the authors chose deliberately because it is the most commonly available spectral configuration in operational coastal mapping, especially for aerial imagery where extra bands are not always accessible. That, combined with publicly released code and pretrained weights and demonstrated performance on Sentinel-2 (10 m) down to 0.25 m aerial imagery, targets operational mapping workflows that need wide coverage at manageable cost.
Future Directions
- Broadening sensor and band coverage. The current study uses only RGB bands and performs no sun glint removal; extending to additional spectral bands, glint correction, or other sensors is an open direction.
- Generalization across domains. The paper's own critique of single-task models centers on degradation when generalizing across geographic domains and coarser resolution; testing Seabed-Net on more sites with different water clarity and substrate reflectance would clarify how far the reported resilience extends.
- Resolving the multi-task trade-offs further. The authors identify cross-task interference and feature entanglement between semantic and geometric representations as unresolved problems, and their ablation and cross-modal analyses point toward deeper disentanglement as a next step.
- Extending beyond bathymetry and classification. The framework's dual-branch fusion design suggests other underwater mapping outputs could be added or substituted, though the paper does not state such extensions.
Note on the reported data: the paper's Table 1 lists 35 patches for Agia Napa and 498 for Puck Lagoon, while the experimental setup section states that 1971 annotated patches were available in Puck Lagoon, of which 1575 were used for training and 396 for evaluation, and that 28 of 35 Agia Napa patches were used for training with 7 for evaluation.
Target Audience
Researchers and practitioners in remote sensing, hydrography, and marine geospatial analysis who work on spectrally-derived bathymetry and seabed habitat mapping; deep learning engineers interested in multi-task architectures with attention- and transformer-based cross-task fusion; and coastal management or surveying teams who need cost-effective depth and habitat products from satellite and aerial imagery. A background in convolutional encoder-decoder networks and transformer attention is required to follow the architecture section, though the experimental comparisons are readable without it.
Authors’ abstract
Accurate, detailed, and regularly updated bathymetry, coupled with complex semantic content, is essential for under-mapped shallow-water environments facing increasing climatological and anthropogenic pressures. However, existing approaches that derive either depth or seabed classes from remote sensing imagery treat these tasks in isolation, forfeiting the mutual benefits of their interaction and hindering the broader adoption of deep learning methods. To address these limitations, we introduce Seabed-Net, a unified multi-task framework that simultaneously predicts bathymetry and pixel-based seabed classification from remote sensing imagery of various resolutions. Seabed-Net employs dual-branch encoders for bathymetry estimation and pixel-based seabed classification, integrates cross-task features via an Attention Feature Fusion module and a windowed Swin-Transformer fusion block, and balances objectives through dynamic task uncertainty weighting. In extensive evaluations at two heterogeneous coastal sites, it consistently outperforms traditional empirical models and traditional machine learning regression methods, achieving up to 75\% lower RMSE. It also reduces bathymetric RMSE by 10-30\% compared to state-of-the-art single-task and multi-task baselines and improves seabed classification accuracy up to 8\%. Qualitative analyses further demonstrate enhanced spatial consistency, sharper habitat boundaries, and corrected depth biases in low-contrast regions. These results confirm that jointly modeling depth with both substrate and seabed habitats yields synergistic gains, offering a robust, open solution for integrated shallow-water mapping. Code and pretrained weights are available at https://github.com/pagraf/Seabed-Net.