Research
Measuring the Intrinsic Dimension of Earth Representations
Overview Research area: Representation learning for Earth observation, specifically geographic implicit neural representations (INRs) / location encoders, evaluated through the lens of intrinsic dimen

- arXiv
- 2511.02101
- Published
- 2025-11-03
- Authors
- Arjun Rao, Marc Rußwurm, Konstantin Klemmer, Esther Rolf
AI summary
Overview
Research area: Representation learning for Earth observation, specifically geographic implicit neural representations (INRs) / location encoders, evaluated through the lens of intrinsic dimension (ID).
Technical level: Intermediate. The paper assumes familiarity with embedding spaces, nearest-neighbor methods, and standard downstream evaluation, but the core idea (how many degrees of freedom a representation actually uses) is accessible.
Scope: A single empirical study that measures the intrinsic dimension of pre-trained geographic INRs across seven location encoders and roughly a dozen geospatial image encoders, linking ID to downstream task performance, spatial resolution, and input modalities.
Author affiliations: University of Colorado Boulder (Rao, Rolf), University of Bonn and Wageningen University (Rußwurm), LGND AI and University College London (Klemmer). Code is released at https://github.com/arjunarao619/GeoINRID.
What This Paper Is About
Geographic INRs take just two numbers (longitude and latitude) and output a high-dimensional embedding vector — typically 256 or 512 dimensions — trained on geo-referenced satellite imagery or text. The problem is that nobody knows how much real information those embeddings contain, or where on the globe that information is concentrated, because evaluation has relied almost entirely on supervised downstream task scores. This paper measures the intrinsic dimension of those embeddings — the number of degrees of freedom actually needed to describe their local variability — as a task-agnostic, label-free way to quantify information content.
Key Contributions
-
First measurement of intrinsic dimension for geographic INRs. The authors provide the first study of ID in implicit neural representations, with a particular focus on geographic ones, measuring both global ID (a single scalar per model) and local ID (a per-location map across Earth's landmass).
-
A two-stage ID framework separating representativeness from task-alignment. ID is computed on frozen pre-trained embeddings to measure representativeness, and on the penultimate-layer activations of a trained downstream model to measure task-alignment.
-
Empirical links between ID and model design choices. The paper shows global ID increases with the spatial resolution hyperparameters of location encoders and with the number of pre-training input modalities, using SatCLIP, GeoCLIP, Sphere2Vec, and Space2Vec as testbeds.
-
Demonstration that local ID maps diagnose spatial artifacts. Local ID estimates expose coverage bias from pre-training data (GeoCLIP), periodic grid patterns from positional encoding design (CSP), and oscillatory artifacts from finite-order spherical harmonics (SatCLIP).
Main Findings
-
Intrinsic dimension is far below ambient dimension but well above 2. For INRs with ambient dimension between 256 and 512, intrinsic dimensions fall roughly between 2 and 10. In Table 1 (distance-based estimators, k=20), SatCLIP–L10 (D=256) records MLE 1.96, MOM 2.02, TLE 2.16, while GeoCLIP (D=512) records MLE 11.21, MOM 13.02, TLE 11.53.
-
Estimator choice changes the number substantially. Angle-based FisherS gives 8.08 for SatCLIP–L40 versus 2.03–2.32 for the distance-based estimators, but drops below 2 for CSP models (1.70 for CSP–fMoW, 0.92 for CSP–iNat).
-
Global ID is stable as embedding size grows, unlike PCA or ICA. In Appendix Table 2(a), SatCLIP's MLE ID stays near 1.96 across D = 64 to 512 while FisherS stays at 5.00; RCF's MLE moves from 5.56 to 6.32. In Table 2(b), PCA at 99% variance retains 36 components at D=64 and 59 at D=512, and ICA retains 64 and 512 respectively, while ID and downstream performance stay roughly flat.
-
Location encoders are competitive with large image encoders. GeoCLIP's ID of 11–13 approaches DOFA (14–16) and CROMA (17–20) on S2-100K Sentinel-2 tiles, and exceeds AlphaEarth (4–9), SINR, and TaxaBind-Sat.
-
Local ID maps reveal where artifacts live. GeoCLIP's local ID is highest in the United States and western Europe, tracing its social-media image pre-training distribution. CSP shows a grid pattern from positional encoding repetition at regular longitude/latitude steps. SatCLIP shows no regional coverage bias but exhibits thin periodic oscillations from finite-order spherical harmonics.
-
Higher global ID in embedding space correlates with better downstream performance. Across air temperature, elevation, population, biome, and countries tasks, higher FisherS global ID of frozen embeddings accompanies higher test R² and top-1 accuracy.
-
Lower global ID in activation space correlates with better downstream performance. Using TwoNN on the penultimate ReLU activations of a 3-hidden-layer MLP yields a strong negative correlation between ID and performance, consistent with task-aligned compression.
-
The same negative relationship appears under direct supervision. Sphere2Vec and Space2Vec encoders trained end-to-end on five TorchSpatial tasks show lower FisherS ID accompanying higher R².
-
More resolution raises global ID. Increasing SatCLIP's Legendre polynomials (L = 10, 20, 40), GeoCLIP's maximum RFF frequency and hierarchy depth M, and Space2Vec/Sphere2Vec's frequency components S all increase global ID. Increasing GeoCLIP's maximum RFF frequency produces a sharp rise; increasing M produces more gradual increases.
-
More input modalities raise both ID and performance. SatCLIP pre-trained on Sentinel-2 only, Sentinel-1 + Sentinel-2, and all MMEarth rasters show increasing FisherS global ID alongside increasing mean test R² on air temperature, elevation, and population density.
Methodology in Plain English
An intrinsic dimension estimator asks a simple question at each point in an embedding space: as you expand a small ball around that point, how quickly does the number of neighbors inside grow? On a curved surface with d degrees of freedom, neighbors accumulate at a characteristic rate, so the growth rate reveals d. The authors use two families of estimators. Distance-based ones (MLE, MOM, TLE) read this from the Euclidean distances to the k nearest neighbors; the paper uses k=20 for the headline Table 1 results and k=100 for local maps. Angle-based FisherS instead re-centers and whitens the embeddings, projects them onto the unit sphere, and reads the effective dimension from how often pairs of directions are nearly parallel — a method robust to local spatial variation, which is why the authors use it for global comparisons on the sphere.
The workflow has two branches. For representativeness, they take a frozen location encoder, feed it 100,000 uniformly sampled land coordinates, and compute global and local ID on the output embeddings. For task-alignment, they train a shallow task head (a 2- or 3-layer MLP; for SustainBench image-location regression a 2-layer MLP branch for image features plus a 5-layer MLP for coordinates) and compute TwoNN ID on the penultimate ReLU activations, evaluated at the dataset's own coordinates. This follows the practice introduced by Ansuini et al. (2019).
Downstream tasks include air temperature, elevation, population density, nightlights, and tree cover regression, plus biome and countries classification and SustainBench image-location regression tasks using TorchSpatial's setup and precomputed InceptionV3 image features. Training runs 20–50 epochs with early stopping, uses Optuna grid search for learning rate, hidden dimension, and weight decay, and reports mean metrics across 10 random seeds. Experiments ran on an NVIDIA Grace-Hopper (GH200) node on CU Boulder's Alpine system.
Why This Matters
Evaluation of Earth representation models has depended on labeling data and training task-specific heads, which is expensive and tells you nothing about a model's general-purpose information content. ID offers a cheap, label-free, architecture-agnostic number that correlates with downstream performance and can be mapped spatially to expose where a model is strong or blind.
Real-world applications suggested or demonstrated:
- Label-free model selection. Comparing architectures, positional encodings, resolution hyperparameters, or input modality sets before any labels are collected.
- Early stopping and training diagnostics. Monitoring divergence between encoder ID and head ID as a signal complementary to validation loss.
- Auditing geographic bias. Local ID maps reveal regional coverage gaps from pre-training data, guiding targeted data collection and region-aware fine-tuning.
- Pre-training design. The modality results give a concrete recipe: adding Sentinel-1 and other MMEarth rasters measurably increases both ID and downstream accuracy.
Industry relevance: organizations building or procuring geospatial foundation models (satellite imagery providers, location-intelligence platforms, environmental monitoring agencies) can use these measures as a pre-screening tool before committing to expensive fine-tuning campaigns. Because the metric needs no labels, it fits naturally into the pre-training stage of large geospatial systems.
Future Directions
-
Extension beyond regression and classification benchmarks. The authors suggest applying ID-based diagnostics to geo-prior settings such as location-conditioned satellite image generation and fine-grained classification with geospatial priors.
-
Attributing information content to data provenance. Complementary measures could credit information content to specific pre-training corpora or regions, clarifying which data sources matter.
-
Localizing information to embedding subspaces. The paper proposes using ID for interpretability by identifying which subspaces hold which content.
-
Using local ID patterns to drive adaptation. More principled ways to let regional variation in local ID steer task-specific fine-tuning or data acquisition strategies.
-
Choosing among estimators. Since FisherS and distance-based estimators disagree substantially across models (falling below 2 for CSP under FisherS but 3–6 under distance-based methods), principled guidance on estimator selection for geographic data remains open.
Target Audience
Researchers and engineers working on geospatial representation learning, remote sensing foundation models, and location encoders. It is also useful for machine learning practitioners interested in intrinsic dimension as an evaluation tool, since the two-stage framework (ID in embedding space versus activation space) generalizes beyond geography. Readers who need a first-principles derivation of the estimators will want to consult the appendix formulas and the cited estimator papers; readers who want the empirical protocol and results can follow the main text.
Authors’ abstract
Within the context of representation learning for Earth observation, geographic Implicit Neural Representations (INRs) embed low-dimensional location inputs (longitude, latitude) into high-dimensional embeddings, through models trained on geo-referenced satellite, image or text data. Despite the common aim of geographic INRs to distill Earth's data into compact, learning-friendly representations, we lack an understanding of how much information is contained in these Earth representations, and where that information is concentrated. The intrinsic dimension of a dataset measures the number of degrees of freedom required to capture its local variability, regardless of the ambient high-dimensional space in which it is embedded. This work provides the first study of the intrinsic dimensionality of geographic INRs. Analyzing INRs with ambient dimension between 256 and 512, we find that their intrinsic dimensions fall roughly between 2 and 10 and are sensitive to changing spatial resolution and input modalities during INR pre-training. Furthermore, we show that the intrinsic dimension of a geographic INR correlates with downstream task performance and can capture spatial artifacts, facilitating model evaluation and diagnostics. More broadly, our work offers an architecture-agnostic, label-free metric of information content that can enable unsupervised evaluation, model selection, and pre-training design across INRs.