Research
AIFloodSense: A Global Aerial Imagery Dataset for Semantic Segmentation and Understanding of Flooded Environments
Overview Research area: Computer vision and remote sensing for disaster management, specifically dataset construction and benchmarking for aerial flood imagery analysis. Technical level: Intermediate.

- arXiv
- 2512.17432
- Published
- 2025-12-19
- Authors
- Georgios Simantiris, Konstantinos Bacharidis, Apostolos Papanikolaou, Petros Giannakakis, Costas Panagiotakis
AI summary
Overview
Research area: Computer vision and remote sensing for disaster management, specifically dataset construction and benchmarking for aerial flood imagery analysis.
Technical level: Intermediate. The paper is a dataset and benchmark paper; understanding it requires familiarity with semantic segmentation, image classification, and visual question answering (VQA), but the core ideas are explained without deep architectural detail.
Scope (one sentence): The paper introduces AIFloodSense, a publicly available dataset of 470 aerial images from 230 flood events across 64 countries and six continents (2022–2024), annotated for classification, pixel-level semantic segmentation, and visual question answering, together with baseline benchmarks.
What This Paper Is About
Floods are becoming more frequent and severe, and aerial imagery is valuable for mapping their extent, but machine learning for this task depends on high-quality annotated datasets that are scarce, geographically narrow, and inconsistent in annotation detail. The authors address this by assembling a globally diverse, temporally recent aerial flood dataset with pixel-level masks for flood, sky, and buildings, plus classification labels for environment type, camera angle, and continent, and a natural-language VQA task. They then establish baseline results for all three task families to define a reproducible benchmark.
Key Contributions
- A new flood segmentation dataset: 470 high-resolution aerial images with pixel-level annotations for flood, buildings, and sky (with a background class), captioned in Table 1 as 4 segmentation classes, providing ancillary information about the built environment and atmospheric visibility.
- Environmental classification annotations: Labels distinguishing rural from urban/peri-urban scenes, presence versus absence of sky in the frame, continent of origin, and latitude/longitude; Table 1 lists 10 classification classes.
- Global and recent coverage: Imagery from flood events in 2022–2024 across all six inhabited continents and 64 countries, spanning 230 distinct flood events, with temporal relevance to contemporary disasters.
- Baseline benchmarks across three tasks: Segmentation baselines over the annotated classes, classification baselines for rural vs. urban/peri-urban and sky presence/absence, a novel continent-prediction benchmark, and baseline VQA results using state-of-the-art methods.
Main Findings
- Scale and diversity: AIFloodSense contains 470 images documenting 230 distinct flood events across 64 countries and six continents (Africa, Asia, Europe, North America, Oceania, South America). Antarctica was excluded due to its geomorphology, lack of permanent habitation, and minimal infrastructure.
- Temporal relevance: Images were captured predominantly in 2023 and 2024, with a small subset from 2022. The yearly breakdown is 5 images from 2022, 209 from 2023, and 256 from 2024.
- Per-continent distribution: Asia 105 images, North America 104, Europe 100, Oceania 60, Africa 51, South America 50.
- Environment and camera-angle balance: 308 images are urban/peri-urban and 162 are rural; 205 show sky presence and 265 show sky absence.
- Resolution: Images were standardized to 1024×768 pixels; the average original resolution is reported as approximately 1.9 megapixels (1,908,570 pixels).
- Annotation quality control: Three trained annotators produced initial masks in Label Studio following detailed guidelines, and two senior supervisors reviewed every image, with iterative refinement until approval. Initial annotation took roughly 50 minutes per image on average, and supervisor-annotator disagreement declined as the process progressed.
- VQA subset construction: Five questions were posed: how many buildings are present, whether the scene is rural or urban/peri-urban, whether sky is visible, whether the image is flooded, and how many buildings are flooded. A Vision-Language Model (Llama3.2) produced initial answers, which were then human-refined.
- Ambiguity filtering for VQA: Because building counts became unreliable with structures at varying depths, a threshold of 80 buildings was applied, filtering out images whose estimated count exceeded it and leaving 294 images. After a second human annotation round using the VLM responses as reference, the final VQA subset is 251 training and 43 test images.
- Split protocol: An 80/20 stratified split balances environment type, camera angle, geographic region, and semantic class pixel proportions between training and test sets.
- Comparison with prior datasets: The related-work comparison (Table 1) positions AIFloodSense as the only listed dataset combining global continental coverage (6), a recent flood date range (2022–2024), aerial imagery, and support for segmentation, classification, and VQA simultaneously. Larger image-count datasets exist (for example FloodNet with 2343 images at 4000×3000 and 10 segmentation classes, and RescueNet with 4494 images at 3000×4000 and 11 segmentation classes), but these are geographically constrained to post-event analysis in Texas, Louisiana, and Florida.
- Baselines: The paper states that baseline benchmarks were established for all tasks using state-of-the-art architectures and that these demonstrate the dataset's complexity, but the specific numerical results are not included in the truncated content provided here.
Methodology in Plain English
The authors sourced high-fidelity aerial images of floods from the World Wide Web, focusing on imagery primarily captured by UAVs to get a bird's-eye view suitable for large-scale assessment. They filtered out blurred or low-quality samples, then annotated each retained image at the pixel level into four semantic classes: flood, sky, building, and background. Every image also received scene-level labels: environment type (rural vs. urban/peri-urban), camera angle (sky present vs. absent), continent, plus metadata such as event date, geolocation, and source URLs.
Annotation ran as a two-stage pipeline: three trained annotators produced initial labels using Label Studio under detailed guidelines covering edge cases like reflections, shadows, and horizon ambiguities; two senior supervisors then reviewed the masks, flagged inconsistencies, and gave feedback until approval. Only approved labels entered the dataset.
For VQA, they first asked five flood-relevant questions to a Vision-Language Model. Because the model miscounted buildings when structures appeared at different depths in a scene — and because two independent human annotators disagreed substantially on such images — they removed images exceeding an estimated 80-building threshold, then had humans re-annotate the remaining 294 images with the VLM answers available as a reference. This yielded a final 251/43 train/test VQA split.
Finally, images were resized to 1024×768 and split 80/20 with stratification so that environment, camera angle, geography, and class pixel proportions stayed balanced. Baseline models were then trained and evaluated on all three task families.
Why This Matters
Impact on research: The paper directly targets a documented gap — the lack of recent, globally distributed flood datasets with pixel-level annotations. Its explicit goal of domain generalization (via continent labels, environment labels, and camera-angle labels) gives researchers a way to measure whether models transfer across regions rather than only within a single flood event. The inclusion of a sky class also addresses a known failure mode where sky is confused with water in datasets lacking that distinction.
Real-world applications:
- Rapid post-disaster damage assessment, including estimating flood extent and counting flooded buildings for emergency response and resource allocation.
- Early warning systems and risk assessment in flood-prone regions, supported by geolocation and event-date metadata.
- Infrastructure and urban planning, using building and flood masks to identify at-risk structures.
- Automated aerial survey pipelines for UAV or satellite operators needing scene triage (rural vs. urban, sky visible or not) before detailed analysis.
- Natural-language interfaces for disaster analysts, where VQA allows non-specialists to query imagery directly.
Industry relevance: Remote sensing providers, insurance and reinsurance firms assessing flood exposure, humanitarian and civil-protection organizations, and companies building UAV-based monitoring products all benefit from a shared, openly available benchmark. Because the paper deliberately excludes synthetic augmentation and pseudo-annotation strategies in order to compare against real-world datasets fairly, baseline numbers are intended to reflect realistic deployment conditions.
Future Directions
- Extending the dataset beyond 470 images and 230 events, and beyond the 2022–2024 window, to cover older and future flood events for longitudinal evaluation.
- Adding more semantic classes and richer annotation granularity, since prior large datasets such as FloodNet and RescueNet label 10 and 11 classes respectively while AIFloodSense labels 4.
- Improving VQA annotation reliability — the authors document systematic VLM miscounting of buildings at varying scene depths and inter-annotator disagreement, which motivates better counting protocols and ambiguity-resolving methods.
- Testing whether the classification priors (continent, environment, camera angle) actually improve segmentation generalization under domain shift, which the paper frames as a motivation but the available content does not quantify.
- Applying or fine-tuning flood-specialized and foundation models on the benchmark, since the baseline evaluation intentionally used general-purpose models rather than disaster-specialized ones.
Target Audience
Researchers and practitioners in computer vision, remote sensing, and disaster management who need a reproducible benchmark for flood segmentation, flood-related classification, or imagery-based question answering. It is also relevant to graduate students and engineers looking for a multi-task aerial dataset, to teams building UAV or satellite flood-monitoring systems, and to humanitarian or government analysts interested in how AI models perform across geographically diverse flood scenes.
Authors’ abstract
Accurate flood detection from visual data is a critical step toward improving disaster response and risk assessment, yet datasets for flood segmentation remain scarce due to the challenges of collecting and annotating large-scale imagery. Existing resources are often limited in geographic scope and annotation detail, hindering the development of robust, generalized computer vision methods. To bridge this gap, we introduce AIFloodSense, a comprehensive, publicly available aerial imagery dataset comprising 470 high-resolution images from 230 distinct flood events across 64 countries and six continents. Unlike prior benchmarks, AIFloodSense ensures global diversity and temporal relevance (2022-2024), supporting three complementary tasks: (i) Image Classification with novel sub-tasks for environment type, camera angle, and continent recognition; (ii) Semantic Segmentation providing precise pixel-level masks for flood, sky, and buildings; and (iii) Visual Question Answering (VQA) to enable natural language reasoning for disaster assessment. We establish baseline benchmarks for all tasks using state-of-the-art architectures, demonstrating the dataset's complexity and its value in advancing domain-generalized AI tools for climate resilience.