Skip to content
AI.info

Research

MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding

Overview Research area: Computer vision and multimodal machine learning for geo-spatial understanding, specifically cross-view image retrieval and geolocalization benchmarks. Technical level: Intermed

arXiv
2512.17492
Published
2025-12-19
Authors
Oskar Kristoffersen, Alba Reinders Sánchez, Morten Rieger Hannemose, Anders Bjorholm Dahl, Dim P. Papadopoulos

AI summary

Overview

Research area: Computer vision and multimodal machine learning for geo-spatial understanding, specifically cross-view image retrieval and geolocalization benchmarks.

Technical level: Intermediate. The paper assumes familiarity with contrastive representation learning (CLIP-style training, InfoNCE loss) and retrieval metrics, but the dataset construction and benchmark setup are described in accessible terms.

Scope: The paper introduces MMLandmarks, a four-modality, instance-level dataset of 18,557 US landmarks, and benchmarks both off-the-shelf foundation models and a simple multimodal baseline (MMCLIP) across cross-view retrieval, geolocalization, and text-to-image/GPS retrieval.

What This Paper Is About

Geo-spatial data about a place can come in many forms: ground photos, aerial imagery, text descriptions, and GPS coordinates. Existing benchmarks typically cover only one or two of these modalities, which encourages specialized models that cannot transfer across tasks. The paper builds a single dataset where every landmark has a one-to-one correspondence across all four modalities, and then measures whether current models can actually exploit that multimodal structure.

Key Contributions

  1. The MMLandmarks dataset. A new dataset with complete pairwise correspondence across four modalities (ground-view images, aerial imagery, text, and GPS coordinates), containing 329,349 ground images, 197,205 aerial images, 18,557 GPS coordinates, and 18,557 textual corpora for 18,557 landmarks in the United States.

  2. A unified multi-task framework. Unlike specialized geo-spatial datasets, MMLandmarks supports training and evaluation on cross-view Ground-to-Satellite retrieval, ground and satellite geolocalization, Text-to-Image, and Text-to-GPS retrieval using the same instance-level annotations.

  3. A benchmark evaluation suite. The paper reports results from off-the-shelf foundation models (DINO, DINOv2, DINOv3, SigLIP, SigLIP2, OAI-CLIP) and task-specific models (Uni-1652, TransGeo-90° FoV, Sample4Geo-UNI, GeoCLIP, StreetCLIP, SatCLIP, G3, GeoReasoner, osv5m) across these tasks.

  4. A versatile baseline (MMCLIP). A simple CLIP-inspired model trained with a complete pairwise contrastive loss over all four modalities, which generalizes across tasks rather than specializing in one.

Main Findings

  • Off-the-shelf foundation models adapt poorly. Zero-shot DINO, SigLIP, SigLIP2, and OAI-CLIP produce very high median ranks on cross-view retrieval — for example, DINO (ViT-B, 224) yields a median rank of 84050 for Satellite-to-Ground retrieval with mAP@1k of 0.3, and DINOv2 (ViT-L, 518) yields 28342. Larger architectures and higher input resolution consistently improve results, with SigLIP2 (ViT-L, 512) attaining the most promising zero-shot numbers (Satellite-to-Ground medR 2461, mAP@1k 1.8; Ground-to-Satellite medR 431, mAP@1k 11.7).

  • Specialized cross-view models transfer badly to MMLandmarks. Uni-1652 gives Satellite-to-Ground medR 62606 and mAP@1k 0.1; TransGeo-90° FoV gives medR 40973 and mAP@1k 0.1; Sample4Geo-UNI gives medR 34988 and mAP@1k 0.4. The authors attribute this to the low image diversity of previous cross-view datasets, which have become saturated.

  • MMCLIP is strong on retrieval but does not saturate the benchmark. On Satellite-to-Ground retrieval it reaches medR 23, mAP@1k 18.8, R@1 30.4, R@5 52.4, and R@10 61.3. On Ground-to-Satellite retrieval it reaches medR 48, mAP@1k 26.2, R@1 20.5, R@5 46.5, and R@10 58.4.

  • Geolocalization shows the same pattern. For Ground-to-GPS, MMCLIP reaches 16.83% at 1 km, 35.95% at 25 km, 51.78% at 200 km, 74.94% at 750 km, and 91.52% at 2500 km, closely tracking GeoCLIP (21.37 / 36.44 / 48.57 / 71.45 / 91.50) and G3 (12.86 / 30.46 / 45.52 / 69.36 / 91.83), while the "Always NYC" baseline achieves 0.00% at 1 km. For Satellite-to-GPS, MMCLIP leads clearly with 36.9% at 1 km, 61.5% at 25 km, 81.1% at 200 km, 95.5% at 750 km, and 99.7% at 2500 km, compared with GeoCLIP (12.3 / 31.3 / 48.8 / 81.3 / 97.4) and G3 (8.8 / 26.9 / 47.9 / 79.8 / 97.0). SatCLIP adapts poorly (0.0 / 0.0 / 0.3 / 5.1 / 50.0), which the authors attribute to the domain gap between Sentinel-2 and NAIP imagery.

  • Text-only queries are harder. For Text-to-Satellite, MMCLIP achieves medR 388, mAP@1k 17.3, R@1 13.4, R@5 31.0, and R@10 41.4, outperforming OAI-CLIP ViT-B/224 (medR 2449, mAP@1k 8.1), OAI-CLIP ViT-L/336 (medR 1037, mAP@1k 14.5), and SigLIP2 ViT-L/512 (medR 3295, mAP@1k 7.16). For Text-to-GPS, MMCLIP reaches 4.2% at 1 km, 16.5% at 25 km, 29.9% at 200 km, 55.2% at 750 km, and 86.1% at 2500 km.

  • Ablations favor fewer modalities, recent imagery, and outdoor-only ground views. Adding modalities slightly decreases retrieval performance: the G,S configuration gives mAP@1k of 17.59 (S→G) and 25.59 (G→S), while the full G,S,T,C configuration gives 17.29 and 26.16. Using the latest satellite image and training on the outdoor-only ground subset improves results; the final MMCLIP configuration reaches 18.79 (S→G), 26.20 (G→S), and 16.83 / 36.9 for Ground/Satellite-to-GPS. An ImageBind-style (G⇔all) objective performs slightly better on retrieval (18.89 and 27.46) but notably worse on Satellite-to-GPS (18.3).

  • Data quality filtering. A LLaVA model (llava-hf/llava-1.5-7b-hf) classified images as indoor or outdoor, yielding 259,451 outdoor and 51,210 indoor images, i.e. outdoor images make up 83% of the original ground views. Manual review of 1000 sampled images found 8.2% wrongly labeled.

Methodology in Plain English

The authors start from OpenStreetMap polygons in the United States that carry a wikipedia or wikidata tag. Each polygon is linked through its Wiki-identifier (format "Q123456") to a Wikipedia page (text) and a Wikimedia Commons page (ground images). To keep landmark sizes comparable, they retain only landmarks whose bounding box's longest edge is under 400 meters. This yields 18,557 landmarks with all four modalities.

Ground images come from Wikimedia Commons and are resized so the largest side is at most 800 pixels. Aerial images come from the NAIP (National Agriculture Imagery Program) via Google Earth Engine, sampled as 800×800 crops at one-to-two-meter pixel size across multiple available years, which introduces natural variation in time, weather, angle, sun orientation, and vegetation.

For evaluation, they build index sets: a ground index derived from GLDv2's 762k images of 101k landmarks, filtered to 17,804 US landmarks, with the 5,277 landmarks overlapping MMLandmarks removed, leaving 714,554 images; and a 100k-image aerial index sampled from noisy coordinate offsets that are more than 500 meters from any training or index coordinate. The query set is 1000 randomly sampled landmarks, contributing 18,688 ground queries and 1000 latest aerial views.

The baseline, MMCLIP, uses a frozen CLIP image encoder for ground and aerial images, the corresponding frozen text encoder for Wikipedia text, and a location encoder initialized following GeoCLIP. Each encoder is followed by a projection head of two trainable linear layers separated by a ReLU. Training extends the InfoNCE loss to all pairwise combinations of the four modalities (K=4), averaged as 1/(K(K−1)) times the sum of pairwise losses. Training uses one NVIDIA H100 GPU, 20 epochs, AdamW with learning rate 1×10⁻⁴, weight decay 5×10⁻⁴, batch size 512, projection dimension 512, and temperature 0.07, with a 1024-landmark validation set.

Why This Matters

Impact on research. The paper argues that geo-spatial learning is inherently multimodal, and that the field's reliance on road-based panoramas and licensing-restricted imagery has produced saturated benchmarks and models that do not generalize. MMLandmarks is presented as the first dataset with complete one-to-one landmark-level correspondence across four modalities, and the first fine-grained, continental-scale cross-view dataset under permissive licenses (Creative Commons/Public Domain for ground imagery, public domain for NAIP), which allows data and models to be shared.

Real-world applications:

  • Fine-grained visual localization for photo organization, geotagging, and digital forensics.
  • Augmented reality and navigation systems that must align ground-level views with overhead maps.
  • Disaster response and infrastructure monitoring, where comparing aerial imagery across timestamps (the NAIP record spans up to two decades in the dataset) supports change detection.
  • Multimodal search over geographic content, matching free-text descriptions to imagery and coordinates.

Industry relevance. Mapping, autonomous navigation, real-estate, insurance, and geospatial-analytics companies all depend on linking ground imagery to overhead imagery and coordinates. The paper's finding that general-purpose foundation models underperform zero-shot, while a comparatively simple model trained on aligned multimodal data generalizes across tasks, is directly relevant to teams deciding whether to fine-tune foundation models or train task-specific baselines.

Future Directions

  • Improving performance on the benchmark, which the authors state their baseline is "far from saturating," particularly for Text-to-Satellite and Text-to-GPS retrieval.
  • Investigating why adding modalities (text and GPS coordinates) slightly reduces cross-view retrieval performance in the ablations, and how to combine modalities without that trade-off.
  • Extending coverage beyond the United States, which the authors describe as a design constraint imposed by NAIP being the only openly available high-resolution aerial source with sufficient geographic diversity.
  • Using the multi-timestamp NAIP imagery for temporal change detection, which the authors flag as an opportunity the dataset opens but do not pursue in this work.

Target Audience

Researchers and practitioners working on cross-view geo-localization, visual place recognition, landmark retrieval, multimodal representation learning, and remote sensing, as well as engineers building location-aware or map-based products who need a benchmark covering multiple geo-spatial tasks under permissive licensing.

Authors’ abstract

Geo-spatial analysis of our world benefits from a multimodal approach, as every single geographic location can be described in numerous ways (images from various viewpoints, textual descriptions, geographic coordinates, etc.). Current benchmarks have limited coverage across modalities, leading to specialized models that perform well in their respective domains, but do not fully take advantage of other geo-spatial modalities. We introduce the Multi-Modal Landmark dataset (MMLandmarks), a benchmark composed of four modalities: 197k high-resolution aerial images, 329k ground-view images, textual information, and geographic coordinates for 18.557 distinct landmarks in the United States. The MMLandmarks dataset has a one-to-one landmark level correspondence across every modality, which enables training and benchmarking models for various geo-spatial tasks, including cross-view Ground-to-Satellite retrieval, ground and satellite geolocalization, Text-to-Image, and Text-to-GPS retrieval. We show that current specialized and off-the-shelf foundation models cannot be trivially used to solve this variety of geo-spatial tasks, illustrating a gap where multimodal datasets lead to broader geo-spatial understanding. We employ a simple CLIP-inspired baseline that reflects versatility and broad generalization when trained with MMLandmarks.

Read the original paper