Research
Few-Shot Adaptation Benchmark for Remote Sensing Vision-Language Models
Overview Research area: Remote sensing (RS) scene classification, vision-language models (VLMs), and few-shot adaptation / benchmarking. Technical level: Intermediate — the reader benefits from famili

- arXiv
- 2510.07135
- Published
- 2025-10-08
- Authors
- Karim El Khoury, Maxime Zanella, Christophe De Vleeschouwer, Benoit Macq
AI summary
Overview
- Research area: Remote sensing (RS) scene classification, vision-language models (VLMs), and few-shot adaptation / benchmarking.
- Technical level: Intermediate — the reader benefits from familiarity with CLIP-style contrastive pretraining and with parameter-efficient adaptation terminology (prompt tuning, adapters, low-rank fine-tuning), though the paper's framing is comparative rather than mathematical.
- Scope: The paper builds and reports the first structured benchmark that evaluates five few-shot adaptation methods applied to three remote sensing vision-language models (plus the original CLIP) across ten RS scene classification datasets.
What This Paper Is About
Remote sensing vision-language models such as RemoteCLIP, GeoRSCLIP, and SkyCLIP achieve strong zero-shot scene classification, but how well they adapt when only a handful of labeled examples per class are available has been largely unexplored. Earlier few-shot adaptation methods (CoOp, MaPLe, TaskRes, Tip-Adapter, CLIP-LoRA) were validated mainly on natural-image CLIP benchmarks, not on RS tasks. This paper supplies a reproducible benchmark that measures those methods on RS data and identifies which models and strategies actually work under low-supervision conditions.
Key Contributions
- First structured few-shot benchmark for RSVLMs. The authors run extensive experiments across ten RS scene classification datasets, three specialized RS vision-language models (plus the original CLIP), and five state-of-the-art few-shot adaptation techniques.
- Evidence that zero-shot accuracy is not predictive of few-shot performance. The study identifies which models and adaptation strategies are best suited to few-shot RS scenarios, showing that models with similar zero-shot accuracy can diverge sharply after adaptation.
- A modular, extensible, open-source codebase for systematic evaluation of RSVLMs under few-shot conditions, released at https://github.com/elkhouryk/fewshot_RSVLMs.
Main Findings
- GeoRSCLIP is the strongest few-shot model. Using a common ViT-B/32 backbone to isolate the effect of pretraining, GeoRSCLIP consistently outperformed RemoteCLIP, SkyCLIP, and CLIP across all five adaptation methods.
- Zero-shot rank does not transfer to few-shot rank. GeoRSCLIP and SkyCLIP show similar zero-shot accuracy on average, yet GeoRSCLIP adapts better; SkyCLIP slightly beats RemoteCLIP zero-shot, but their few-shot performance is nearly equivalent across values of K.
- One-shot adaptation already beats zero-shot. Even a single labeled example per class produced significant improvements over zero-shot for all four models considered.
- No single adaptation method wins everywhere. CLIP-LoRA has the best average accuracy at 2-shot and 4-shot on GeoRSCLIP ViT-B/32 (84.4 and 87.8 average, versus CoOp's 83.0 and 87.1), while TaskRes is best at 1-shot (79.7 average) and excels on MLRSNet and RESISC45, the datasets with the highest number of classes.
- Tip-Adapter reverses its ranking as shots increase. It is the weakest method at 1-shot (71.3 average on GeoRSCLIP ViT-B/32) but becomes the strongest at 8 and 16 shots (91.5 and 94.0 average), surpassing CoOp (90.2 / 92.6), CLIP-LoRA (91.2 / 93.4), MaPLe (87.2 / 90.4), and TaskRes (86.5 / 90.0).
- MaPLe performs worst overall, which the authors attribute to its design focus on generalizing to new classes and/or domains.
- Accuracy and training cost trade off. CLIP-LoRA and CoOp require gradient-based optimization and cost more training time, while embedding-space methods like TaskRes offer a better balance in the 1- and 2-shot settings. CLIP-LoRA retains a practical advantage: its low-rank matrices can be merged into model weights after training, adding no inference cost. Tip-Adapter's cache grows with shots, increasing inference time.
- Low-rank fine-tuning is most robust to backbone scaling. CLIP-LoRA leads on ViT-L/14 (307M parameters) and ViT-H/14 (986M parameters), while CoOp, MaPLe, and TaskRes drop significantly on ViT-H/14 — MaPLe falling from 79.5 average at 2-shot on ViT-B/32 to 69.4 on ViT-H/14.
Methodology in Plain English
The authors assemble ten RS scene classification datasets — AID, EuroSAT, MLRSNet, OPTIMAL31, PatternNet, RESISC45, RSC11, RSICB128, RSICB256, and WHURS19 — ranging from roughly 10^3 to 10^5 samples and from 10 to 46 classes. All samples are grouped and split 50%/25%/25% into train, validation, and test sets with a fixed random seed, and none of the datasets were used during pretraining of the models.
They then take three RS-specific VLMs (RemoteCLIP, GeoRSCLIP, SkyCLIP) plus the original CLIP, each with its published visual and textual backbones, and prompt them with RS-specific templates of the form "a satellite photo of a [class]." For each downstream task they build a support set of K labeled examples per class, with K tested at 1, 2, 4, 8, and 16, and evaluate on a disjoint query set of the same classes.
Five adaptation methods are applied using each method's original published hyperparameters: CoOp (prompt tuning over learnable text tokens), MaPLe (learnable visual tokens added to textual ones), TaskRes (task-specific residual tuning of class text embeddings), Tip-Adapter (a cache combining stored features with zero-shot predictions), and CLIP-LoRA (low-rank matrices inserted into text and visual encoders). For the model comparison, all models are run on the same ViT-B/32 backbone to isolate pretraining effects from architecture, and results are averaged over three random seeds. Backbone scaling is then studied on GeoRSCLIP's ViT-B/32, ViT-L/14, and ViT-H/14 configurations, and training time is measured on an NVIDIA A100 80GB GPU.
Why This Matters
Impact on research. The paper establishes that strong zero-shot numbers are not a proxy for few-shot quality in remote sensing, so future RSVLM work must report adaptation results rather than zero-shot scores alone. The absence of a clear winner among the five methods — and the fact that the best method changes with the number of shots, the dataset, and the backbone size — frames few-shot adaptation for RS as an open problem rather than a solved recipe. The released codebase and fixed splits make comparisons reproducible and extensible.
Real-world applications (the paper motivates RS scene classification with these settings):
- Environmental monitoring, where analysts classify large volumes of imagery with limited labeled examples.
- Data-driven farming, where crop and land-cover categories must be recognized quickly from satellite imagery.
- Emergency disaster response, where rapid scene understanding is needed and labeled data for a new event is scarce by definition.
- General RS scene classification pipelines where annotation is expensive relative to the scale of available imagery.
Industry relevance. The benchmark's cost analysis matters for deployment: CLIP-LoRA's mergeable low-rank matrices avoid added inference cost, while cache-based Tip-Adapter grows its inference burden with more shots. Because accuracy, training time, and inference time trade off against each other, the paper argues that practitioners should weigh computational constraints alongside accuracy. The authors do not report an industrial deployment; the relevance is that the benchmark and its open-source framework let teams choose an adaptation strategy fit to their supervision level and compute budget.
Future Directions
- Develop few-shot methods tailored to remote sensing, rather than importing methods designed for natural-image CLIP.
- Build methods that stay strong across supervision levels, since the leading method shifts from TaskRes or CLIP-LoRA at 1 to 4 shots to Tip-Adapter at 8 and 16 shots.
- Report computation alongside accuracy as standard practice, because the most accurate method is not always the most practical one.
- Extend the benchmark beyond scene classification to other RS tasks, and systematically evaluate across backbone sizes, since results on ViT-B/32 do not carry over unchanged to ViT-L/14 or ViT-H/14.
Target Audience
Researchers and engineers working on remote sensing, geospatial machine learning, or vision-language models who need to adapt a pretrained model with very few labels. It is also useful for practitioners selecting between prompt tuning, adapter-based, and low-rank fine-tuning strategies under real compute constraints, and for benchmark builders who want a reproducible template for few-shot evaluation in a specialized domain.
Authors’ abstract
Remote Sensing Vision-Language Models (RSVLMs) have shown remarkable potential thanks to large-scale pretraining, achieving strong zero-shot performance on various tasks. However, their ability to generalize in low-data regimes, such as few-shot learning, remains insufficiently explored. In this work, we present the first structured benchmark for evaluating few-shot adaptation methods on RSVLMs. We conduct comprehensive experiments across ten remote sensing scene classification datasets, applying five widely used few-shot adaptation strategies to three state-of-the-art RSVLMs with varying backbones. Our findings reveal that models with similar zero-shot performance can exhibit markedly different behavior under few-shot adaptation, with some RSVLMs being inherently more amenable to such adaptation than others. The variability of performance and the absence of a clear winner among existing methods highlight the need for the development of more robust methods for few-shot adaptation tailored to RS. To facilitate future research, we provide a reproducible benchmarking framework and open-source code to systematically evaluate RSVLMs under few-shot conditions. The source code is publicly available on Github: https://github.com/elkhouryk/fewshot_RSVLMs