This paper presents Text-Guided Backdoor (TGB), an adjustable backdoor attack against multimodal pretrained models that uses natural-word triggers, namely words that can naturally occur in ordinary textual inputs. Most existing backdoor attacks require specific trigger conditions that are typically not satisfied by ordinary inference inputs, thereby limiting their activation in real-world deployments. TGB overcomes this limitation by exploiting naturally occurring words as triggers, enabling stealthy activation without requiring explicit trigger insertion at inference time. This property avoids conspicuous trigger patterns and improves the practicality of TGB. Furthermore, we introduce visual adversarial perturbations on poisoned samples to modulate the model's learning of natural-word triggers, thereby enabling flexible adjustment of TGB attack strength without modifying the poisoned data. Extensive experiments are conducted on downstream tasks built upon multimodal pretrained models, including Composed Image Retrieval (CIR) and Visual Question Answering (VQA). Results demonstrate the effectiveness of TGB across diverse poisoning settings and its ability to flexibly adjust attack success rates, which reveal critical security vulnerabilities in multimodal pretrained models.
Adjustable Text-Guided Backdoor Attacks with Natural-Word Triggers on Multimodal Pretrained Models
This paper presents Text-Guided Backdoor (TGB), an adjustable backdoor attack against multimodal pretrained models that uses natural-word triggers, namely words that can naturally occur in ordinary textual inputs.
- Preview

- Year
- 2026
- Hosting
- Abstract onlyARXIV-DEFAULT
Cite
Notes
Only stored in your browser.
Attribution
- Abstract & full text
- arxiv.org/abs/2604.05809ARXIV-DEFAULT
- TL;DR
- Semantic Scholar