A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models
PHANTOM is an open dataset of ~50k multimodal jailbreaks for evaluating the safety of Vision-Language Models. It offers 47,524 adversarial image-text pairs spanning 7,826 harmful intents across 10 categories and 55 subcategories, built with five attack strategies in single- and multi-turn settings. Attacks are generated against six open-source VLMs and transferred to six proprietary frontier models along with nine open-source VLMs.
01 โ Dataset
We evaluate six open-source VLMs (DeepSeek-VL2, GLM-4.6V-Flash, Kimi-VL-A3B-Instruct, Qwen3-VL-30B-A3B-Instruct, Qwen3.5-27B, Qwen3.6-27B) and transfer the resulting attacks to six proprietary frontier models (GPT-5.4, GPT-5.5, Claude Opus 4.6/4.7/4.8, Gemini 3.1 Pro), reporting the Attack Success Rate (ASR, %) per model, per harmful-intent category, and per attack strategy. Higher ASR indicates a more vulnerable model.
Bold red values mark the row-wise (per model) maximum across categories, including Average. Click any category header (or Average) to re-rank the table by that column, ascending. Source: PHANTOM benchmark corpus, ASR (%).
Category legend
02 โ Benchmark
Harmful intents are curated and stratified into ten top-level categories โ including a dedicated Child Safety category โ and further split into 55 fine-grained subcategories, giving broad coverage of the risk surface relevant to multimodal jailbreaks. Each intent is paired with one or more attacks built using BAP, IDEATOR, MML, FC_ATTACK and CSDJ. See the leaderboard for per-category attack success rates.
| Category | # Intents |
|---|---|
| A. Ethical and Social Risks | 988 |
| B. Privacy and Data Risks | 504 |
| C. Safety and Physical Harm | 877 |
| D. Criminal and Economic Risks | 1,017 |
| E. Cybersecurity Threats | 725 |
| F. Information and Political Manipulation | 534 |
| G. Content and Cultural Safety | 537 |
| H. Intellectual Property and Ownership | 304 |
| I. Decision and Cognitive Risks | 1,593 |
| J. Child Safety (new) | 747 |
TABLE โ Number of intents per category
03 โ Analysis
We further analyze failure modes across attack strategies, intent categories, and model families, examining where transfer from open-source surrogate models to proprietary targets succeeds or fails, and which visual-textual perturbation patterns most reliably bypass current safety alignment.
04 โ Citation
@misc{gallivanone2026phantomlargescaledatasetmultimodal,
title={PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models},
author={Simone Gallivanone and Hossein Khodadadi and Mauro Dore and Mauro Medda and Nicola Franco},
year={2026},
eprint={2606.24388},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2606.24388}
}