PHANTOM

A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Simone Gallivanone1,*, Hossein Khodadadi1,*, Mauro Dore2, Mauro Medda2, Nicola Franco1

* Equal contribution

1 The Italian Institute of Artificial Intelligence (AI4I), Turin, Italy
2 HikmaAI S.r.l., Pula, Italy

PHANTOM teaser figure
Fig. 1 โ€” Overview of the PHANTOM pipeline: harmful intents, attack strategies, target models, and evaluation.

PHANTOM is an open dataset of ~50k multimodal jailbreaks for evaluating the safety of Vision-Language Models. It offers 47,524 adversarial image-text pairs spanning 7,826 harmful intents across 10 categories and 55 subcategories, built with five attack strategies in single- and multi-turn settings. Attacks are generated against six open-source VLMs and transferred to six proprietary frontier models along with nine open-source VLMs.

47,524Adversarial pairs
7,826Harmful intents
10 / 55Categories / subcategories
5Attack strategies
9 / 6Open / proprietary VLMs

01 โ€” Dataset

PHANTOM Dataset

We evaluate six open-source VLMs (DeepSeek-VL2, GLM-4.6V-Flash, Kimi-VL-A3B-Instruct, Qwen3-VL-30B-A3B-Instruct, Qwen3.5-27B, Qwen3.6-27B) and transfer the resulting attacks to six proprietary frontier models (GPT-5.4, GPT-5.5, Claude Opus 4.6/4.7/4.8, Gemini 3.1 Pro), reporting the Attack Success Rate (ASR, %) per model, per harmful-intent category, and per attack strategy. Higher ASR indicates a more vulnerable model.

Bold red values mark the row-wise (per model) maximum across categories, including Average. Click any category header (or Average) to re-rank the table by that column, ascending. Source: PHANTOM benchmark corpus, ASR (%).

Category legend

AEthical and Social Risks BPrivacy and Data Risks CSafety and Physical Harm DCriminal and Economic Risks ECybersecurity Threats FInformation and Political Manipulation GContent and Cultural Safety HIntellectual Property and Ownership IDecision and Cognitive Risks JChild Safety
Benchmark overview plot
Fig. 2 โ€” Average ASR across attack strategies, by model.
Attack statistics plot
Fig. 3 โ€” Statistics of generated attacks by category and model.

02 โ€” Benchmark

PHANTOM Benchmark

Harmful intents are curated and stratified into ten top-level categories โ€” including a dedicated Child Safety category โ€” and further split into 55 fine-grained subcategories, giving broad coverage of the risk surface relevant to multimodal jailbreaks. Each intent is paired with one or more attacks built using BAP, IDEATOR, MML, FC_ATTACK and CSDJ. See the leaderboard for per-category attack success rates.

Category# Intents
A. Ethical and Social Risks988
B. Privacy and Data Risks504
C. Safety and Physical Harm877
D. Criminal and Economic Risks1,017
E. Cybersecurity Threats725
F. Information and Political Manipulation534
G. Content and Cultural Safety537
H. Intellectual Property and Ownership304
I. Decision and Cognitive Risks1,593
J. Child Safety (new)747

TABLE โ€” Number of intents per category

Category coverage radar plot
Fig. 4 โ€” Distribution of harmful intents across risk categories.

03 โ€” Analysis

PHANTOM Attack Analysis

We further analyze failure modes across attack strategies, intent categories, and model families, examining where transfer from open-source surrogate models to proprietary targets succeeds or fails, and which visual-textual perturbation patterns most reliably bypass current safety alignment.

Attack analysis figure
Fig. 5 โ€” Cross-model and cross-category attack transfer analysis.

04 โ€” Citation

BibTeX

CITE THIS WORK
@misc{gallivanone2026phantomlargescaledatasetmultimodal,
  title={PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models},
  author={Simone Gallivanone and Hossein Khodadadi and Mauro Dore and Mauro Medda and Nicola Franco},
  year={2026},
  eprint={2606.24388},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
  url={https://arxiv.org/abs/2606.24388}
}