Research
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Overview Research area: Cross-modal (video-text) retrieval in computer vision, applied to a new domain — drone or UAV fo

- arXiv
- 2510.15470
- Published
- 2025-10-17
- Authors
- Jinghao Huang, Yaxiong Chen, Ganchao Liu
AI summary
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text RetrievalOverview
- Research area: Cross-modal (video-text) retrieval in computer vision, applied to a new domain — drone or UAV footage. The paper sits at the intersection of vision-language representation learning and aerial imagery analysis.
- Technical level: Intermediate. The method builds on CLIP encoders and contrastive retrieval, with probabilistic embeddings added on top; readers should be comfortable with attention, contrastive losses, and the standard video-text retrieval evaluation protocol.
- Scope: The paper proposes a method (MSAM) and two datasets (USRD, UMCRD) for the task of retrieving drone video given text and retrieving text given drone video.
What This Paper Is About
Drone cameras produce enormous quantities of overhead video, but existing video-text retrieval models were designed for ground-level footage and do not handle the repetitive structure, high visual similarity, and varied wording found in aerial scenes. The authors propose drone video-text retrieval (DVTR) as a task, argue it needs dedicated mechanisms, and build MSAM to model the many possible semantic descriptions that a single drone video can have. They also construct two drone video-text datasets to serve as benchmarks.
Key Contributions
- A multi-semantic adaptive learning mechanism that adaptively generates multi-semantic embeddings for each drone video and text, promoting semantic consistency alignment, enhancing cross-modal information flow, and improving the model's diversity in semantic learning.
- A cross-modal interactive feature fusion pooling (CIFFP) mechanism that lets the model infer which drone video frames are most relevant to a given text, guiding it toward the core semantic content of the description and optimizing the supervision signal.
- Two new drone video-text datasets — the UAV Scene Retrieval Dataset (USRD) and the UAV Multi-Context Retrieval Dataset (UMCRD) — described as the first of their kind in this field and intended as standard evaluation benchmarks.
- The claim of being the first to systematically propose and study the drone video-text retrieval task, with experiments on both self-constructed datasets showing MSAM outperforming competing methods.
Main Findings
- New task framing: Drone video differs from common video in content and shooting perspective — common video focuses on human behavior and temporal alignment, while drone video emphasizes scene structure and target objects. Overhead views with repetitive roads, buildings, and vehicles create high visual similarity across samples, which the authors argue intensifies cross-modal ambiguity.
- UMCRD results (Text → Video): MSAM reports R@1 of 49.5, R@5 of 84.0, R@10 of 92.2, MdR of 2, and MnR of 4.8. The strongest listed comparison is CLIP4Clip at R@1 48.0, R@5 82.9, R@10 91.9, MdR 2, MnR 5.2.
- UMCRD results (Video → Text): MSAM reports R@1 of 55.7, R@5 of 91.7, R@10 of 97.8, MdR of 1, and MnR of 2.5. The strongest listed comparison is CLIP4Clip at R@1 54.3, R@5 90.0, R@10 97.0, MdR 1, MnR 2.9.
- USRD results (Text → Video): MSAM reports R@1 of 29.9, R@5 of 71.2, R@10 of 90.3, MdR of 3, and MnR of 5.9. The closest listed comparison is TS2-Net at R@1 29.3 and UCoFiA at R@1 29.2.
- USRD results (Video → Text): MSAM reports R@1 of 33.8, R@5 of 81.3, R@10 of 94.8, MdR of 2, and MnR of 4.2. The closest listed comparisons are X-Pool at R@1 32.8 and TempMe at R@1 32.5.
- Comparison breadth: Baselines span DRL, Frozen, CLIP4Clip, X-Pool, CenterClip, X-Clip, TS2-Net, UATVR, UCoFiA, T-MASS, DGL, GLSCL, and TempMe (published between 2021 and 2025). MSAM exceeds all of them on the reported R@1 figures for both directions on both datasets.
- Dataset statistics: USRD contains 2,864 videos of 5s, 1,043 words, average caption length 10.32, average caption similarity 0.026, and dataset diversity 3.43. UMCRD contains 4,235 videos of 5s to approximately 8s, 2,275 words, average caption length 18.33, average caption similarity 0.032, and dataset diversity 4.86.
- Dataset composition: UMCRD comprises 4,235 samples, of which 2,372 were collected on YouTube and 1,863 were filmed manually with drones. USRD's 2,864 samples come from the ERA dataset, with each drone video sized 640 × 640 pixels. Each drone video is paired with five text descriptions.
- Most frequent content: Person, road, building, and plant are the most frequent object names; green and red are the most common color descriptors; "middle" and "surround" are the most common spatial descriptions. Annotations come from manual labeling and from the Gemini API, with prompts inspired by ShareGPT4V.
- Ablation (partial, as presented): On UMCRD Text → Video, CLIP4Clip reaches R@1 48.0, R@5 82.9, R@10 91.9, MdR 2, MnR 5.2; adding CIFFP gives 49.1 / 83.4 / 92.0 / 2 / 5.0; adding L_ddsl gives 49.8 / 84.1 / 92.1 / 2 / 4.9. On USRD Text → Video the same progression is 28.6 / 69.7 / 89.7 / 3 / 6.2, then 29.5 / 70.1 / 89.9 / 3 / 6.1, then 29.7 / 70.9 / 90.3 / 2 / 5.9. The remaining ablation rows and the parameter-tuning results are not included in the provided excerpt.
- Data splits: USRD is split 70% / 20% / 10% into training, validation, and test sets; UMCRD is split 60% / 20% / 20%.
Methodology in Plain English
The system starts from CLIP. A text encoder turns each sentence into a feature vector plus a sequence of token features, and a visual encoder does the same for each frame of a drone video, producing a video feature and a per-frame feature set.
The first component, CIFFP, addresses the fact that in drone footage target objects can be occluded, disappear, or reappear, and ordinary pooling ignores the text entirely. CIFFP computes similarity between every frame embedding and the text embedding, converts those scores into attention weights with softmax, and uses them to aggregate the frames into a text-aware video representation. That representation is then compared against the text to score text relevance, and a learned sigmoid gate decides, per video, how much to weight video information versus text information. A residual connection adds the fused result back onto the original aggregated video features, and a final matrix multiplication produces the video-text similarity score used for ranking.
The second component, multi-semantic adaptive learning, handles the one-to-many nature of drone video description. Instead of a single deterministic embedding per sample, the model produces k probabilistic embeddings for each video and each text. Attention over the token and frame features produces intermediate representations, which pass through LayerNorm and a fully connected layer, and then through two heads predicting a mean vector and a variance for each of the k embeddings. Variance across the k means represents model uncertainty; the mean of the k variances represents data uncertainty.
Training uses three losses. The video-text matching term is a standard bidirectional contrastive loss over the similarity scores, treating matched pairs in a batch as positives and all others as negatives, with a temperature parameter. The distribution-driven semantic learning term directly minimizes the gap between the video and text Gaussian distributions by comparing their means and variances, rather than relying only on contrastive loss plus Monte Carlo estimation. The diversity semantic term applies orthogonal constraints so that the k embeddings capture different semantics instead of collapsing into correlated features — it pushes the Gram matrix of the k embeddings toward the identity matrix under the Frobenius norm. The total objective is the video-text matching term plus the distribution-driven term plus a weighted diversity term, with a trade-off parameter λ.
Why This Matters
- Impact on research: The paper stakes out drone video-text retrieval as a distinct problem rather than a domain transfer of ground-level retrieval, and releases two datasets with a quantitative diversity analysis. It also connects probabilistic embedding ideas previously used for image-text retrieval to the aerial video setting, and argues that prior probabilistic methods only replicate many-to-many correspondences without properly constraining the distributions.
- Real-world applications (as listed in the paper): urban planning; agricultural monitoring; natural disaster monitoring and rescue; urban traffic management. The paper also notes drone advantages over satellites — lower cost, real-time high-resolution video, fast decisions from real-time streaming, and no cloud-cover weather limitation.
- Industry relevance: The motivating pressure is data volume — rapidly growing drone video creates an urgent need for efficient semantic retrieval so that operators can find relevant clips by describing them. Search over aerial footage is useful for mapping, inspection, surveillance, and media production pipelines, and the two released datasets give vendors and researchers a common benchmark, with the paper stating that source code and datasets will be made publicly available.
Future Directions
- Scaling the benchmarks: Both datasets are in the low thousands of videos (USRD 2,864; UMCRD 4,235), and the paper pre-emptively argues this is sufficient for robustness. Whether MSAM's advantages hold on much larger, more heterogeneous drone corpora is an open question.
- Real-time deployment: The paper highlights real-time streaming as a drone advantage over satellites, but the work reports offline retrieval benchmarks. Latency and efficiency of the CIFFP and k-embedding pipeline on live drone feeds are not reported in the available content.
- Interaction with detection and tracking: The authors note that existing drone literature focuses on vehicle or human detection and tracking. Connecting retrieval with target detection or tracking would let a user query a moving object across a long flight rather than a short clip.
- Cross-domain and cross-platform generalization: It is not established whether a model tuned on USRD and UMCRD transfers to other aerial platforms, satellite imagery, or other geographic regions; the hyperparameter λ and the choice of k also leave room for further study, and the parameter-tuning experiments are not included in the excerpt reviewed here.
Target Audience
Researchers and graduate students working on cross-modal retrieval, vision-language models, or aerial/UAV computer vision; practitioners building search over drone footage archives or real-time aerial monitoring systems; and benchmark builders who need drone video-text datasets with documented statistics and splits. Readers new to CLIP-based retrieval will need to follow up on the cited methods (X-Pool, X-CLIP, PCME, UATVR) to follow every design choice.
Authors’ abstract
With the advancement of drone technology, the volume of video data increases rapidly, creating an urgent need for efficient semantic retrieval. We are the first to systematically propose and study the drone video-text retrieval (DVTR) task. Drone videos feature overhead perspectives, strong structural homogeneity, and diverse semantic expressions of target combinations, which challenge existing cross-modal methods designed for ground-level views in effectively modeling their characteristics. Therefore, dedicated retrieval mechanisms tailored for drone scenarios are necessary. To address this issue, we propose a novel approach called Multi-Semantic Adaptive Mining (MSAM). MSAM introduces a multi-semantic adaptive learning mechanism, which incorporates dynamic changes between frames and extracts rich semantic information from specific scene regions, thereby enhancing the deep understanding and reasoning of drone video content. This method relies on fine-grained interactions between words and drone video frames, integrating an adaptive semantic construction module, a distribution-driven semantic learning term and a diversity semantic term to deepen the interaction between text and drone video modalities and improve the robustness of feature representation. To reduce the interference of complex backgrounds in drone videos, we introduce a cross-modal interactive feature fusion pooling mechanism that focuses on feature extraction and matching in target regions, minimizing noise effects. Extensive experiments on two self-constructed drone video-text datasets show that MSAM outperforms other existing methods in the drone video-text retrieval task. The source code and dataset will be made publicly available.