Research
REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent
REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent Overview Research area: Remote sensing (RS) and computer vision, sitting at the intersection of foundation model (FM)

- arXiv
- 2511.17442
- Published
- 2025-11-21
- Authors
- Binger Chen, Tacettin Emre Bök, Behnood Rasti, Volker Markl, Begüm Demir
AI summary
REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware AgentOverview
Research area: Remote sensing (RS) and computer vision, sitting at the intersection of foundation model (FM) benchmarking, structured knowledge databases, and LLM-based tool-using agents.
Technical level: Intermediate. The paper assumes familiarity with foundation models, dense retrieval, retrieval-augmented generation (RAG), and LLM agent orchestration, but it explains each component clearly without requiring deep mathematical background.
Scope: The paper introduces RS-FMD (a structured database of over 160 remote sensing foundation models) and Remsa (a constraint-aware LLM agent that selects suitable RSFMs from free-text queries), together with a 100-query expert-scored benchmark and rubric-based evaluation protocol.
What This Paper Is About
Remote sensing practitioners now have hundreds of foundation models to choose from, spanning different sensor modalities (RGB, multispectral, hyperspectral, SAR, LiDAR, text), pretraining strategies, and downstream tasks. The problem is that picking the right one for a specific task, sensor, region, and compute budget is a manual, time-consuming, and hard-to-reproduce process, because model documentation is scattered across papers, model cards, and repositories with no unified schema. The goal of this paper is to formalize that upstream selection step and automate it with a domain-specific agent grounded in structured model metadata.
Key Contributions
-
RS-FMD, a structured foundation model database. The authors introduce what they describe as the first structured, schema-guided database of over 160 remote sensing foundation models, covering data modalities, architectures, pretraining datasets, reported benchmarks, and deployment properties. It is populated with a semi-automated, schema-guided LLM extraction pipeline with confidence-guided human verification, and will be released as a maintained community resource.
-
Remsa, a modular constraint-aware selection agent. The agent combines structured metadata grounding, dense retrieval, in-context LLM ranking, interactive clarification, explanation generation, and memory augmentation under a task-aware orchestration mechanism. It is designed to be LLM-agnostic and operates entirely on publicly available metadata of open-source RSFMs, without accessing private or sensitive data.
-
A new benchmark and rubric-based expert evaluation protocol. The authors construct 100 realistic natural language RS query scenarios and define a seven-criterion suitability rubric scored by two domain experts on a 1–5 scale, producing 3,000 expert-scored task–system–model configurations across 4 systems and 3 LLM backbones.
-
Carefully constructed baselines for component analysis. Three systems (Remsa-Naive, DB-Retrieval, and Unstructured-RAG) isolate the contribution of orchestration, LLM-based ranking, and structured schema grounding respectively.
Main Findings
-
Remsa leads on every metric under GPT-4.1. Remsa achieves an Average Top-1 Score of 75.76, Average Set Score of 75.03, Top-1 Hit Rate of 21.33%, High-Quality (HQ) Hit Rate of 40.00%, and MRR of 0.34. The next-best system, Remsa-Naive, scores 72.67, 72.00, 20.00%, 37.33%, and 0.29.
-
Structured decision logic beats retrieval alone. Compared to DB-Retrieval (67.37 Avg Top-1, 68.87 Avg Set, 12.00% Top-1 Hit, 17.33% HQ Hit, 0.23 MRR), Remsa raises Top-1 Hit Rate from 12.00% to 21.33%, HQ Hit Rate from 17.33% to 40.00%, and MRR from 0.23 to 0.34 — attributed to composing multiple constraints rather than direct metadata matching.
-
Unstructured RAG underperforms despite full model descriptions. Unstr.-RAG reaches a moderate Average Top-1 of 71.23, but its Top-1 Hit Rate (13.33%) and MRR (0.24) remain substantially below Remsa, suggesting that structured schema grounding and modular reasoning matter more than raw access to text.
-
Orchestration adds measurable value over naive tool use. Both Remsa and Remsa-Naive clearly outperform the retrieval-only and unstructured-RAG baselines, but Remsa improves further over Remsa-Naive on Avg Top-1 (72.67 to 75.76), HQ Hit Rate (37.33% to 40.00%) and MRR (0.29 to 0.34), pointing to the contribution of multi-turn clarification and decision heuristics.
-
Results hold across LLM backbones. With DeepSeek3.2, Remsa records 75.35 Avg Top-1, 73.81 Avg Set, 18.67% Top-1 Hit, 40.00% HQ Hit, and 0.30 MRR. With LLaMA-3.3-70B, it records 73.39, 70.34, 14.67%, 32.00%, and 0.26. Remsa outperforms the corresponding baselines on all three backbones, and HQ Hit Rate is stable across GPT-4.1 and DeepSeek3.2 (both 40.00%).
-
A latency cost accompanies the quality gain. Average end-to-end runtime per query is 0.77s for DB-Retrieval, 11.9s for Unstr.-RAG, 22.7s for Remsa-Naive, and 31.7s for Remsa, which the authors attribute to multi-stage decision processing and optional clarification.
-
Rubric sensitivity analysis (truncated in the provided text). The paper reports that removing criteria changes scores: full scoring yields 75.03 Avg Set, 22.67% Top-1 Hit, and 0.38 MRR; removing Application Compatibility gives 73.32, 21.33%, 0.36; removing Modality Match gives 70.88, 22.67%, 0.36. The table is cut off at the "w/o Reporte[d Performance]" row, so further results are not reported in the provided content.
-
Confidence calibration for database curation. Fields flagged with a confidence score below θ = 0.75 are sent for human review, with the score combining a normalized log-probability term (weight 0.7) and a self-consistency term (weight 0.3), at temperature τ = 0.5. A manual inspection of all fields for 10 records found high-confidence outputs to be consistently accurate.
Methodology in Plain English
The authors attacked the problem in two stages. First, they built the resource that makes selection possible at all: RS-FMD. They searched surveys, venues, arXiv keyword queries, and linked GitHub repositories, then converted the messy text of papers and model cards into machine-readable JSON records following a fixed schema. Rather than extracting fields by hand, they used an LLM to generate each field several times, validated the outputs against the schema, and scored every field by how confident the model was and how consistent its answers were across sampling rounds. Fields that fell below the threshold went to a human for verification, so annotators only looked at the risky parts, not whole records.
Second, they built Remsa on top of that database. A user types a query in ordinary language. An interpreter converts it into structured constraints, extracting the application and the required modality as mandatory fields, and treating data availability, compute budget, fine-tuning needs, and output-quality priorities as optional ones. A task orchestrator then runs a control loop: it checks how complete the constraints are, how large the candidate pool is, whether anything violates a hard requirement, and how confident the ranking is, and then chooses which tool to call. The retrieval tool encodes constraints and database entries with Sentence-BERT embeddings (with type-indicator tokens prefixed to each field) and searches them with FAISS using cosine similarity, optimized for high recall so that soft matches are not lost. The ranking tool then removes candidates that break hard constraints using deterministic rules and re-ranks the rest with an LLM prompted with expert-crafted few-shot examples. If constraints are missing or confidence is low, the clarification tool asks up to three follow-up questions. Finally, the explanation tool produces a structured report with the model name, justification bullets, and links to the paper and code repository.
For evaluation, the authors wrote 100 realistic query scenarios from templates covering tasks such as flood mapping with SAR, crop type classification with multispectral or hyperspectral imagery, urban expansion monitoring with optical time series, sea ice detection, and wildfire detection. Each system returned its top-3 models; those model–query pairs were anonymized and blindly scored by two experts on seven criteria. During evaluation, Remsa's clarification rounds were executed automatically against a separate LLM simulating user responses, so no human was involved and no evaluator bias could creep in across systems.
Why This Matters
Impact on research. The paper reframes foundation model selection as a first-class research problem: an upstream stage that determines what is compared and deployed before any downstream adaptation happens. It provides a machine-readable schema, a reusable database, and a reproducible expert protocol with publicly released queries, guidelines, scoring criteria, and model metadata — all of which other groups can build on rather than recreating ad hoc spreadsheets.
Real-world applications:
- Flood mapping and disaster response using SAR imagery, where the required sensor modality and latency constraints dictate which model is viable.
- Crop type classification with multispectral or hyperspectral imagery, where spectral resolution and regional pretraining coverage matter.
- Urban expansion monitoring with optical time series, requiring multi-temporal inputs and spatial resolution compatible with city-scale analysis.
- Environmental monitoring and hazard detection such as sea ice and wildfire detection, often run under tight compute budgets.
Industry relevance. Organizations deploying remote sensing pipelines need to justify model choices to stakeholders and reproduce them across projects. Remsa produces transparent justifications and links to papers and code repositories, which supports auditability. Because it works from free-text queries and asks clarifying questions, it lowers the barrier for practitioners who are not remote sensing specialists. Its reliance on only public open-source metadata means no private or sensitive data is involved, which removes a common procurement obstacle.
Future Directions
-
Closing the loop with downstream benchmarking. The authors state explicitly that exhaustive adaptation and evaluation of every candidate RSFM under every query-specific setting is infeasible and beyond the scope of this work. Measuring whether rubric-based suitability scores actually predict downstream task accuracy after fine-tuning is the obvious open question.
-
Extending error mitigation and feedback. The paper notes that an optional LLM-as-a-Judge component could re-evaluate low-quality selections, and that the modular design is extensible to more robust error mitigation strategies. Building out those feedback mechanisms is a natural next step.
-
Sustaining and growing the database. RS-FMD is planned for public maintenance with monthly rolling checks by maintainers, community submissions from model authors, and automatic schema extraction plus maintainer review before inclusion. The paper also describes a fallback "closest-match" mode for when no candidate fully satisfies constraints.
-
Generalizing beyond remote sensing. The authors frame their contribution as formalizing model selection under modality, data, compute, and application constraints, and release the evaluation resources as a foundation for future work "in RS and beyond." The confidence weights (0.7 and 0.3) are explicitly stated as adjustable rather than fixed, inviting practitioners to recalibrate them for different LLMs or domains.
Target Audience
This paper is most useful to remote sensing scientists and machine learning researchers who need to choose among foundation models rather than train one; to industry practitioners and engineers building operational RS pipelines who need documented, reproducible, and auditable model selection; to LLM agent researchers interested in a concrete, constraint-heavy domain application rather than general-purpose assistance; and to benchmark designers looking for a worked example of rubric-based expert evaluation for a task where no single ground-truth answer exists.
Authors’ abstract
Foundation Models (FMs) are increasingly integrated into remote sensing (RS) pipelines. These models include unimodal vision encoders and multimodal architectures. FMs are adapted to diverse perception tasks, such as image classification, change detection, and visual question answering. However, selecting the most suitable remote sensing foundation model (RSFM) for a specific task remains challenging due to scattered documentation, heterogeneous formats, and complex deployment constraints. To address this, we first introduce the RSFM Database (RS-FMD), the first structured and schema-guided resource covering over 160 RSFMs trained on various data modalities, spanning different spatial, spectral, and temporal resolutions, considering different learning paradigms. Built upon RS-FMD, we further present REMSA, a constraint-aware agent that enables automated RSFM selection from natural language queries. REMSA combines structured FM metadata retrieval with a task-driven decision workflow. In detail, it interprets user input, clarifies missing constraints, ranks models via in-context learning, and provides transparent justifications. Our system supports various RS tasks and data modalities, enabling personalized, reproducible, and efficient FM selection. To evaluate REMSA, we construct a benchmark of 100 expert-verified RS query scenarios. Each query is evaluated across 4 systems and 3 LLM backbones, with the top-3 selected models manually assessed by domain experts. This results in 3,000 expert-scored task--system--model configurations under our novel expert-centered evaluation protocol. REMSA outperforms multiple baselines, showing its practical utility in real decision-making applications. REMSA operates entirely on publicly available metadata of open source RSFMs, without accessing private or sensitive data.