Research
Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning
Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning Overview Research area: Computer vision security and adversarial robustness, specifically backdoor attacks on open-vo

- arXiv
- 2511.12735
- Published
- 2025-11-16
- Authors
- Ankita Raj, Chetan Arora
AI summary
Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt TuningOverview
Research area: Computer vision security and adversarial robustness, specifically backdoor attacks on open-vocabulary object detectors (OVODs) adapted through parameter-efficient prompt tuning.
Technical level: Advanced. The paper assumes familiarity with vision-language architectures (Grounding DINO, GLIP), prompt tuning methods (CoOp, CoCoOp, VPT), and object detection evaluation metrics.
Scope: The paper proposes TrAP (Trigger-Aware Prompt tuning), the first backdoor attack framework for open-vocabulary object detectors, which jointly optimizes learnable prompts in the image and text branches alongside a visual trigger patch.
What This Paper Is About
Open-vocabulary object detectors such as Grounding DINO and GLIP can detect arbitrary categories from text prompts and are increasingly used in robotics, autonomous driving, and surveillance, but nobody had studied whether they can be secretly backdoored. The authors show that when a pre-trained OVOD is adapted to a downstream dataset using lightweight prompt tuning, an attacker with white-box access during that adaptation can implant a hidden trigger that causes objects to be misclassified or to disappear entirely. The goal is to demonstrate this attack surface and build an attack that is both effective and hard to detect, while actually improving the model's accuracy on clean images.
Key Contributions
-
First study of backdoor attacks on open-vocabulary object detectors. The authors state that no prior work examined backdoor threats in OVODs, and they frame prompt tuning as a new attack surface because it is modular, lightweight, and avoids retraining or duplicating the full model.
-
TrAP: a multi-modal backdoor injection strategy. Learnable prompt tokens are inserted into both the vision branch (following VPT-Deep, with m_v = 50 tokens per layer) and the text branch (a CoCoOp-inspired variant the authors call CoCoOp-Det, with m_t = 4 tokens), while a visual trigger patch is learned jointly with these prompts.
-
A curriculum-based trigger shrinking strategy. Training begins with a larger trigger (ρ = 0.2 of the object bounding box) for the first 10 epochs and reduces it to ρ = 0.1 for the remaining 5 epochs, so that a small, inconspicuous patch activates the backdoor at inference time.
-
Two attack objectives plus extensive evaluation. The paper formalizes Object Misclassification Attack (OMA) and Object Disappearance Attack (ODA), evaluates them across six ODinW-13 datasets on Grounding DINO and GLIP, benchmarks against adapted CoCoOp-Det and VPT baselines, and tests robustness against three inference-time defenses.
Main Findings
-
TrAP outperforms single-modality baselines on OMA (Table 1). On Vehicles, TrAP reaches ASR 0.79 with BmAP 64.87, versus VPT at ASR 0.64 and CoCoOp-Det at ASR 0.08. Across the six datasets TrAP records ASR values of 0.79 (Vehicles), 0.88 (Aquarium), 0.83 (Aerial Drone), 0.75 (Shellfish), 0.92 (Thermal), and 1.00 (Mushrooms).
-
Clean performance improves over zero-shot in nearly every setting. For example, zero-shot mAP on Aerial Drone is 15.1 while TrAP reaches BmAP 46.00; on Thermal, zero-shot is 54.2 versus TrAP BmAP 78.17.
-
ODA results require care in interpretation (Table 2). Most methods report ASR of 1.00, so the authors argue that a successful disappearance attack must show high BAP with low PAP. TrAP keeps BAP near the benign level (e.g., Shellfish BAP 58.37 versus VPT 58.03) while dropping PAP to 6.93 on Shellfish, 3.60 on Aquarium, 6.83 on Vehicles, 24.63 on Thermal, and 26.33 on Mushrooms.
-
Both modalities are needed. On Vehicles, text-only CoCoOp-Det achieves higher benign mAP (66.8) than image-only VPT (64.0), but VPT is far better at learning the trigger association (ASR 0.64 versus 0.08), leading the authors to conclude that tuning both branches balances clean accuracy and attack strength.
-
Curriculum learning matters. Training and testing at ρ = 0.2/0.1 yields ASR 0.66, and training at ρ = 0.1 without curriculum yields ASR 0.75, both below TrAP's 0.79 (Table 3).
-
Prompt tuning is far cheaper than fine-tuning. TrAP trains 0.2M parameters, while Fine-tune-A uses 36M and Fine-tune-B uses 21M parameters; the fine-tuned variants achieve higher BmAP (67.4 and 68.2) but lower ASR (0.74 and 0.67) and the authors note they require training over 100× more parameters.
-
The meta-net contributes to clean accuracy. Removing the instance-specific context drops BmAP from 64.9 to 63.9 on Vehicles.
-
Trigger size trades off against stealth. ρ = 0.5 gives ASR 0.90 and BmAP 65.0, while ρ = 0.05 still gives ASR 0.75 with BmAP 63.3.
-
The attack transfers to GLIP (Table 4). Using GLIP's native text prompting with VPT in the vision branch, TrAP reaches ASR 0.89 on Vehicles (BmAP 62.2), 0.96 on Aquarium, 0.89 on Aerial Drone, 0.84 on Shellfish, 0.96 on Thermal, and 1.00 on Mushrooms.
-
Existing inference-time defenses largely fail (Table 5). PatchDrop at 50% lowers BmAP by about 20 points while ASR stays at 0.63; prompt rewording is inconsistent (Bus→Buses ASR 0.71, Bus→Omnibus 0.49, Bus→A Bus 0.04); the PAD adversarial patch defense on a checkerboard trigger slightly raised ASR from 0.48 to 0.50.
Methodology in Plain English
The researchers do not touch the pre-trained model's original weights. Instead they keep Grounding DINO frozen and add two sets of small trainable "prompt" vectors: one set inserted into every transformer layer of the image encoder, and one set prepended to each class name's word embeddings in the text branch. A small two-layer network (Linear–ReLU–Linear with 16× dimension reduction) generates an image-conditioned vector that is added to the text prompts, so the class representations adapt to what is in the picture. Alongside these prompts, a square trigger patch is itself learned as a trainable variable.
Each training batch is split so that the model is optimized with two losses summed together: a clean loss on unmodified images that keeps detection accurate, and a poisoned loss on images stamped with the trigger whose annotations have been rewritten to the attacker's target (either a wrong class label for OMA, or removal of the target class for ODA). A weight λ = 1 balances the two. The trigger patch is placed at the center of the relevant bounding boxes at a size that is a fraction ρ of the box, and the curriculum progressively shrinks ρ from 0.2 to 0.1 so that small patches remain effective at test time.
Training uses 15 epochs with AdamW at learning rate 0.001, batch size 4, on a single NVIDIA Tesla V100 32GB GPU. Evaluation uses standard COCO mAP at IoU thresholds 0.5 to 0.95, reporting benign mAP (BmAP), poisoned mAP (PmAP), and ASR defined as triggered boxes with confidence above 0.5 and IoU above 0.5 predicted as the target class divided by the total non-target boxes; ODA analogously reports benign AP (BAP), poisoned AP (PAP), and ASR as the fraction of triggered target-class boxes that vanish.
Why This Matters
Impact on research. The paper opens an entirely unexamined threat category: security of prompt-tuned open-vocabulary detectors. It also shows that a defense widely used against backdoors in image classifiers (patch dropping) does not transfer cleanly, and that the ASR metric alone is misleading for disappearance attacks, which is a methodological point other security researchers will need to account for. The counterintuitive finding that the poisoned model can be more accurate on clean data than the zero-shot model makes the backdoor harder to flag by ordinary performance monitoring.
Real-world applications at risk:
- Autonomous driving, where a sticker on an ambulance could cause it to be classified as an ordinary vehicle, exactly the scenario the authors sketch, or could cause a vehicle to disappear from perception entirely.
- Surveillance and monitoring systems, where a small trigger patch could suppress detection of a target class.
- Robotics, where trigger-conditioned misclassification could redirect manipulation or navigation decisions.
- Third-party model adaptation services, since the threat model assumes a user outsources fine-tuning of a private dataset (a few hundred annotated examples) to an external party with white-box access.
Industry relevance. The attack is cheap to mount: 0.2M trainable parameters and no access to the massive pre-training corpora of Grounding DINO or GLIP. For companies that routinely fine-tune open-vocabulary detectors via cloud services or third-party libraries, this represents a supply-chain risk that standard deployment practices do not currently address.
Future Directions
- Developing defenses specifically tailored to OVOD backdoors, which the authors explicitly list as the main limitation of the current work.
- Addressing the overfitting the authors observe on a single category name: prompt engineering with "A Bus" reduced ASR to 0.04, and the authors hypothesize this can be alleviated by training on variations of category names, leaving it for future work.
- Extending the framework to additional attack objectives; the authors state in the supplementary material that the work can be extended to Object Generation Attacks.
- Broadening evaluation and threat scenarios, since the study covers six of the ODinW-13 datasets (six single-category datasets and PascalVOC were excluded) and two victim models.
Target Audience
Researchers and practitioners in computer vision security, adversarial machine learning, and trustworthy AI; engineers deploying open-vocabulary detectors in safety-critical products; and anyone using prompt tuning or third-party adaptation services for large vision-language models. Readers need a working understanding of transformer architectures, prompt tuning, and detection metrics, though the threat model and attack mechanics are explained well enough for a security-focused reader without deep detection expertise.
Authors’ abstract
Open-vocabulary object detectors (OVODs) unify vision and language to detect arbitrary object categories based on text prompts, enabling strong zero-shot generalization to novel concepts. As these models gain traction in high-stakes applications such as robotics, autonomous driving, and surveillance, understanding their security risks becomes crucial. In this work, we conduct the first study of backdoor attacks on OVODs and reveal a new attack surface introduced by prompt tuning. We propose TrAP (Trigger-Aware Prompt tuning), a multi-modal backdoor injection strategy that jointly optimizes prompt parameters in both image and text modalities along with visual triggers. TrAP enables the attacker to implant malicious behavior using lightweight, learnable prompt tokens without retraining the base model weights, thus preserving generalization while embedding a hidden backdoor. We adopt a curriculum-based training strategy that progressively shrinks the trigger size, enabling effective backdoor activation using small trigger patches at inference. Experiments across multiple datasets show that TrAP achieves high attack success rates for both object misclassification and object disappearance attacks, while also improving clean image performance on downstream datasets compared to the zero-shot setting. Code: https://github.com/rajankita/TrAP