liuxuannan/Awesome-Multimodal-Jailbreak

A Survey on Jailbreak Attacks and Defenses against Multimodal Generative Models

335

403 commits

updated Jan 11, 2026

See the code

README

😈🛡️Awesome-Jailbreak-against-Multimodal-Generative-Models

🔥🔥🔥 Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey

Paper

We've curated a collection of the latest 😋, most comprehensive 😎, and most valuable 🤩 resources on Jailbreak Attack and Defense against Multimodel Generative Models.
But we don't stop there; Our repository is constantly updated to ensure you have the most current information at your fingertips.

survey model

🤗Introduction

This survey presents a comprehensive review of existing jailbreak attack and defense against multimodal generative models.
Given the generalized lifecycle of multimodal jailbreak, we systematically explore attacks and corresponding defense strategies across four levels: input, encoder, generator, and output.

🧑‍💻 Four Levels of Multimodal Jailbreak lifecycle

  • Input Level: Attackers and defenders operate solely on the input data. Attackers modify inputs to execute attacks, while defenders incorporate protective cues to enhance detection.
  • Encoder Level: With access to the encoder, attackers optimize adversarial inputs to inject malicious information into the encoding process, while defenders work to prevent harmful information from being encoded within the latent space.
  • Generator Level: : With full access to the generative models, attackers leverage inference information, such as activations and gradients, and fine-tune models to increase adversarial effectiveness, while defenders use these techniques to strengthen model robustness.
  • Output Level: With the output from the generative model, attackers can iteratively refine adversarial inputs, while defenders can apply post-processing techniques to enhance detection.

Based on this analysis, we present a detailed taxonomy of attack methods, defense mechanisms, and evaluation frameworks specific to multimodal generative models.
We cover a wide range of input-output configurations, including modalities such as Any-to-Text, Any-to-Vision, and Any-to-Any within generative systems.

survey model

🚀Table of Contents

🔥Multimodal Generative Models

Below are tables of model short name and representative generative models used for jailbreak. For input/output modalities, I: Image, T: Text, V: Video, A: Audio.

📑Any-to-Text Models (LLM Backbone)

Short NameModalityRepresentative Model
I+T→TI + T → TLLaVA, MiniGPT4, InstructBLIP
VT2TV + T → TVideo-LLaVA, Video-LLaMA
AT2TA + T → TAudio Flamingo, Audiopalm

📖Any-to-Vision (Diffusion Backbone)

Short NameModalityRepresentative Model
T→IT → IStable Diffusion, Midjourney, DALLE
IT→II + T → IDreamBooth, InstructP2P
T2VT → VOpen-Sora, Stable Video Diffusion
IT2VI + T → VVideoPoet, CogVideoX

📰Any-to-Any (Unified Backbone)

Short NameModalityRepresentative Model
IT→ITI + T → I + TNext-GPT, Chameleon
TIV2TIVT + I + V → T + I + VEMU3
Any2AnyAny → AnyGPT-4o, Gemini Ultra

😈JailBreak Attack

📖Attack-Intro

We categorize attack methods into black-box, gray-box, and white-box attacks. in a black-box setting where the model is inaccessible to the attacker, the attack is limited to surface-level interactions, focusing solely on the model’s input and/or output. Regarding gray-box and white-box attacks, we consider model-level attacks, including attacks at both the encoder and generator.

  • Input-level attack: attackers are compelled to develop more sophisticated input templates across prompt engineering, image engineering, and role-ploy techniques.
  • Output-level attack: Attackers focus on querying outputs across multiple input variants. Driven by specific adversarial goals, attackers employ estimation-based and search-based attack techniques to iteratively refine these input variants.
jailbreak_attack_black_box
  • Encoder-level attack: Attackers are restricted to accessing only the encoders to provoke harmful responses. In this case, attackers typically seek to maximize cosine similarity within the latent space, ensuring the adversarial input retains similar semantics to the target malicious content while still being classified as safe.
  • Generator-level attack: Attackers have unrestricted access to the generative model’s architecture and checkpoint, enabling attackers to conduct thorough investigations and manipulations, thus enabling sophisticated attacks.
jailbreak_attack_white_and_gray_box

📑Papers

Below are the papers related to jailbreak attacks.

Jailbreak Attack of Any-to-Text Models

TitleVenueDateCodeTaxonomyMultimodal Model
Jailbreaking Large Vision Language Models in Intelligent Transportation SystemsArxiv 20252025/11/17None---I+T→T
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMsArxiv 20252025/11/20None---I+T→T
The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental JailbreaksArxiv 20252025/11/20None---I+T→T
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMsNeurIPS 20252025/05/17HomepageInput LevelV+T→T
Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-IntensityArxiv 20252025/08/11Github---I+T→T
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual SteeringACM MM 20252025/08/07Github---I+T→T
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM JailbreakingArxiv 20252025/07/29None---I+T→T
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language ModelsArxiv 20252025/07/20None---I+T→T
Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language ModelsTMM 20252025/07/18None---I+T→T
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context InjectionArxiv 20252025/07/03Github---I+T→T
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language ModelsTCSVT 20252025/07/02Github---I+T→T
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language ModelsArxiv 20252025/05/26Github---I+T→T
VSCBench: Bridging the Gap in Vision-Language Model Safety CalibrationArxiv 20252025/05/26Github---I+T→T
JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language ModelsArxiv 20252025/05/26None---I+T→T
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language ModelsArxiv 20252025/05/26None---A+T→T
Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box FrameworArxiv 20252025/05/24None---A+T→T
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language ModelsArxiv 20252025/05/23None---A+T→T
BadNAVer: Exploring Jailbreak Attacks On Vision-and-Language NavigationArxiv 20252025/05/22None---I+T→T
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language ModelsArxiv 20252025/05/20None---A+T→T
Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language ModelsArxiv 20252025/05/18None---I+T→T
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning ModelArxiv 20252025/05/10Github---I+T→T
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning ModelsArxiv 20252025/04/09Github---I+T→T
PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code ContextualizationArxiv 20252025/04/02None---I+T→T
Multilingual and Multi-Accent Jailbreaking of Audio LLMsArxiv 20252025/04/01None---A+T→T
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution StrategyCVPR 20252025/03/26Github---I+T→T
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak AttacksArxiv 20252025/03/24None---I+T→T
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV EqualizationArxiv 20252025/03/14None---I+T→T
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist ContentArxiv 20252025/03/13None---I+T→T
Utilizing Jailbreak Probability to Attack and Safeguard Multimodal LLMsArxiv 20252025/03/10None---I+T→T
FC-Attack: Jailbreaking Large Vision-Language Models via Auto-Generated FlowchartsArxiv 20252025/02/28None---I+T→T
EigenShield: Causal Subspace Filtering via Random Matrix Theory for Adversarially Robust Vision-Language ModelsArxiv 20252025/02/20None---I+T→T
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMsArxiv 20252025/02/16None---I+T→T
Distraction is All You Need for Multimodal Large Language Model JailbreakingCVPR 20252025/02/15None---I+T→T
ELITE: Enhanced Language-Image Toxicity Evaluation for SafetyArxiv 20252025/02/07None---I+T→T
`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMsArxiv 20252025/02/02None---A+T→T
"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language ModelsArxiv 20252025/02/02Github---A+T→T
Failures to Find Transferable Image Jailbreaks Between Vision-Language ModelsICLR 20252025/01/23GithubGenerator LevelI+T→T
Jailbreaking Multimodal Large Language Models via Shuffle InconsistencyICCV 20252025/01/09None---I+T→T
Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language ModelsArxiv 20242024/12/21None---I+T+A→T
AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language ModelsICLR 20252024/12/11GithubGenerator LevelA+T→T
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language ModelsICCV 20252024/12/8Github---I+T→T
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity MaximizationArxiv 20242024/12/8None---I+T→T
Jailbreak Large Vision-Language Models Through Multi-Modal LinkageArxiv 20242024/11/30Github---I+T→T
Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language ModelsCVPR 20252024/11/27None---I+T→T
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and DefenseArxiv 20242024/11/13None---I+T→T
MMJ-Bench : A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language ModelsArxiv 20242024/08/16None---I+T→T
Failures to Find Transferable Image Jailbreaks Between Vision-Language ModelsNeurIPS 2024 Workshops2024/07/21None---I+T→T
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language ModelsNeurIPS 20242024/06/11Github---I+T→T
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak AttacksArxiv 20242024/06/10Github---I+T→T
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?Arxiv 20242024/04/04Github---I+T→T
VLSBench: Unveiling Visual Leakage in Multimodal SafetyACL 20252024/11/29HomepageInput LevelI+T→T
Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language ModelsArxiv 20242024/11/18GithubOutput LevelI+T→T
IDEATOR: Jailbreaking Large Vision-Language Models Using ThemselvesICCV 20252024/11/15GithubOutput LevelI+T→T
Zer0-Jack: A memory-efficient gradient-based jailbreaking method for black box Multi-modal Large Language ModelsNeurIPS SafeGenAi Workshop 20242024/11/12NoneOutput LevelI+T→T
Audio is the achilles’heel: Red teaming audio large multimodal modelsArxiv 20242024/10/31GithubInput LevelA+T→T
Advweb: Controllable black-box attacks on vlm-powered web agentsArxiv 20242024/10/22NoneInput LevelI+T→T
Can Large Language Models Automatically Jailbreak GPT-4V?NAACL Workshop 20242024/07/23NoneInput LevelI+T→T
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak PromptsACM MM 20242024/07/21NoneInput LevelI+T→T
Image-to-Text Logic Jailbreak: Your Imagination can Help You Do AnythingArxiv 20242024/07/01NoneInput LevelI+T→T
From LLMs to MLLMs: Exploring the Landscape of Multimodal JailbreakingEMNLP 20242024/06/21NoneEncoder LevelI+T→T
Jailbreak Vision Language Models via Bi-Modal Adversarial PromptArxiv 20242024/06/06GithubGenerator LevelI+T→T
Efficient LLM-Jailbreaking by Introducing Visual ModalityArxiv 20242024/05/30GithubGenerator LevelI+T→T
White-box Multimodal Jailbreaks Against Large Vision-Language ModelsACM MM 20242024/05/28GithubGenerator LevelI+T→T
Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language ModelsArxiv 20242024/05/26Github---I+T→T
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image CharacterArxiv 20242024/05/25GithubInput LevelI+T→T
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language ModelsECCV 20242024/05/14GithubGenerator LevelI+T→T
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially FastICML 20242024/02/13GithubGenerator LevelI+T→T
Jailbreaking Attack against Multimodal Large Language ModelArxiv 20242024/02/04GithubGenerator LevelI+T→T
Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsICLR 2024 Spotlight2024/01/16GithubEncoder LevelI+T→T
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language ModelsECCV 20242023/11/29GithubInput LevelI+T→T
How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMsECCV 20242023/11/27GithubEncoder LevelI+T→T
Jailbreaking GPT-4V via Self-Adversarial Attacks with System PromptsArxiv 20232023/11/15NoneOutput LevelI+T→T
FigStep: Jailbreaking Large Vision-language Models via Typographic Visual PromptsAAAI 20252023/11/09GithubInput LevelI+T→T
Image Hijacks: Adversarial Images can Control Generative Models at RuntimeICML 20242023/09/01GithubGenerator LevelI+T→T
Are aligned neural networks adversarially aligned?NeurIPS 20232023/06/26NoneGenerator LevelI+T→T
Visual Adversarial Examples Jailbreak Aligned Large Language ModelsAAAI 20242023/06/22GithubGenerator LevelI+T→T
On Evaluating Adversarial Robustness of Large Vision-Language ModelsNeurIPS 20232023/05/26HomepageEncoder LevelI+T→T

Jailbreak Attack of Any-to-Vision Models

TitleVenueDateCodeTaxonomyMultimodal Model
Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model SafeguardsACM MM 20252025/07/30Github---T→I
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image ModelsArxiv 20252025/07/23None---T→I
PLA: Prompt Learning Attack against Text-to-Image Generative ModelsICCV 20252025/07/14None---T→I
GhostPrompt: Jailbreaking Text-to-image Generative Models based on Dynamic OptimizationArxiv 20252025/05/25None---T→I
TokenProber: Jailbreaking Text-to-image Models via Fine-grained Word Impact AnalysisArxiv 20252025/05/11None---T→I
T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak AttacksArxiv 20252025/05/10None---T→V
Inception: Jailbreak the Memory Mechanism of Text-to-Image Generation SystemsArxiv 20252025/04/29None---T→I
Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image ModelsArxiv 20252025/04/15None---T→I
Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive JailbreakingCVPR 2025 Highlight2025/04/08Github---T→I
Reason2Attack: Jailbreaking Text-to-Image Models via LLM ReasoningArxiv 20252025/03/23None---T→I
Jailbreaking Safeguarded Text-to-Image Models via Large Language ModelsArxiv 20252025/03/03None---T→I
Unified Prompt Attack Against Text-to-Image Generation ModelsTPAMI 20252024/02/23None---T→I
T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image GenerationArxiv 20252025/02/22Github---T→I
CogMorph: Cognitive Morphing Attacks for Text-to-Image ModelsArxiv 20252024/01/21None---T→I
FameBias: Embedding Manipulation Bias Attack in Text-to-Image ModelsArxiv 20242024/12/24None---T→I
Antelope: Potent and Concealed Jailbreak Attack StrategyArxiv 20242024/12/11None---T→I
Multimodal Pragmatic Jailbreak on Text-to-image ModelsArxiv 20242024/09/27None---T→I
In-Context Experience Replay Facilitates Safety Red-Teaming of Text-to-Image Diffusion ModelsArxiv 20242024/11/25NoneOutput LevelT→I
Unfiltered and Unseen: Universal Multimodal Jailbreak Attacks on Text-to-Image Model DefensesOpenreview2024/11/13None---T→I
AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion modelsArxiv 20242024/10/28GithubEncoder LevelT→I
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by StepArxiv 20242024/10/4NoneOutput LevelT→I
ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep GenerationNeurIPS 20242024/9/25GithubInput LevelT→I
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image ModelsArxiv 20242024/08/25NoneOutput LevelT→I
Perception-guided Jailbreak against Text-to-Image ModelsAAAI 20252024/08/20GithubInput LevelT→I
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order OptimizationNAACL 20252024/08/18GithubOutput LevelT→I
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion ModelsArxiv 20242024/08/02NoneEncoder LevelT→I
Jailbreaking Text-to-Image Models with LLM-Based AgentsArxiv 20242024/08/01NoneOutput LevelT→I
Automatic Jailbreaking of the Text-to-Image Generative AI SystemsICML 2024 Workshop NextGenAISafety2024/05/26GithubOutput LevelT→I
UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual CheckersICML 20242024/05/18NoneInput LevelT→I
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image GeneratorsArxiv 20242024/02/23NoneInput LevelT→I
Harnessing LLM to Attack LLM-Guarded Text-to-Image ModelsArxiv 20232023/12/12GithubInput LevelT→I
MMA-Diffusion: MultiModal Attack on Diffusion ModelsCVPR 20242023/11/29GithubEncoder LevelT→I
VA3: Virtually Assured Amplification Attack on Probabilistic Copyright Protection for Text-to-Image Generative ModelsCVPR 20242023/11/29GithubGenerator LevelT→I
To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For NowECCV 20242023/10/18GithubGenerator LevelT→I
Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?ICLR 20242023/10/16GithubEncoder LevelT→I
SurrogatePrompt: Bypassing the Safety Filter of Text-To-Image Models via SubstitutionCCS 20242023/09/25NoneInput LevelT→I
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic PromptsICML 20242023/09/12GithubGenerator LevelT→I
SneakyPrompt: Jailbreaking Text-to-image Generative ModelsSymposium on Security and Privacy 20242023/05/20GithubOutput LevelT→I
Red-Teaming the Stable Diffusion Safety FilterNeurIPSW 20222022/10/03NoneInput LevelT→I

Jailbreak Attack of Any-to-Any Models

TitleVenueDateCodeTaxonomyMultimodal Model
Gradient-based Jailbreak Images for Multimodal Fusion ModelsArxiv 20242024/10/4GithubGenerator LevelI+T→I+T
Voice jailbreak attacks against gpt-4oArxiv 20242024/05/29GithubOutput LevelAny→Any

🛡️Jailbreak Defense

📖Defense-Intro

Current efforts made in the jailbreak defense of multimodal generative models include two lines of work: Discriminative defense and Transformative defense.

  • Discriminative defenses: is constrained to classification tasks for assigning binary labels.
jailbreak_discriminative_defense
  • Transformative Defense: aims to produce appropriate and safe responses in the presence of malicious or adversarial inputs.
jailbreak_transformative_defense

📑Papers

Below are the papers related to jailbreak defense.

Jailbreak Defense of Any-to-Text Models

TitleVenueDateCodeTaxonomyMultimodal Model
Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model SecurityArxiv 20252025/11/20Github---I+T→T
Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models: A Unified and Accurate ApproachArxiv 20252025/08/08None---I+T→T
Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model SecurityArxiv 20252025/07/29None---I+T→T
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore MechanismArxiv 20252025/07/02None---I+T→T
The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language ModelsArxiv 20252025/06/15None---I+T→T
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language ModelsArxiv 20252025/05/28None---I+T→T
GuardReasoner-VL: Safeguarding VLMs via Reinforced ReasoningArxiv 20252025/05/16Github---I+T→T
DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language ModelsArxiv 20252025/04/25Github---I+T→T
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?CVPR 20252025/04/14None---I+T→T
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language ModelArxiv 20252025/04/3Github---I+T→T
Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based AttacksICCV 20252025/04/2Github---I+T→T
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial DefenseArxiv 20252025/03/14None---I+T→T
Utilizing Jailbreak Probability to Attack and Safeguard Multimodal LLMsArxiv 20252025/03/10None---I+T→T
Adversarial Training for Multimodal Large Language Models against Jailbreak AttacksArxiv 20252025/03/05None---I+T→T
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden StatesACL 20252025/02/20Github---I+T→T
SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine UnlearningArxiv 20252025/02/18None---I+T→T
Understanding and Rectifying Safety Perception Distortion in VLMsArxiv 20252025/02/18None---I+T→T
Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial TrainingArxiv 20252025/02/17None---I+T→T
Towards Robust Multimodal Large Language Models Against Jailbreak AttacksArxiv 20252025/02/02None---I+T→T
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language ModelsArxiv 20252025/01/30Github---I+T→T
Internal Activation Revision: Safeguarding Vision Language Models Without Parameter UpdateArxiv 20252025/01/24None---I+T→T
MSTS: A Multimodal Safety Test Suite for Vision-Language ModelsArxiv 20252025/01/17Github---I+T→T
Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language ModelsArxiv 20252025/01/03Github---I+T→T
Defending LVLMs Against Vision Attacks through Partial-Perception SupervisionArxiv 20242024/12/17None---I+T→T
VLMGuard: Defending VLMs against Malicious Prompts via Unlabeled DataArxiv 20242024/10/01None---I+T→T
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time AlignmentCVPR 20252024/11/27GithubOutput LevelI+T→T
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against JailbreaksCVPR 20252024/11/23GithubGenerator LevelI+T→T
Uniguard: Towards universal safety guardrails for jailbreak attacks on multimodal large language modelsArxiv 20242024/11/03NoneInput LevelI+T→T
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single VectorArxiv 20242024/10/30GithubGenerator LevelI+T→T
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak AttacksICLR 20252024/10/28GithubInput LevelI+T→T
The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?Arxiv 20242024/10/02NoneInput LevelI+T→T
CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional CalibrationCOLM 20242024/9/17NoneOutput LevelI+T→T
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial AttacksBigData 20242024/09/11NoneEncoder LevelI+T→T
Bathe: Defense against the jailbreak attack in multimodal large language models by treating harmful instruction as backdoor triggerArxiv 20242024/08/17NoneGenerator LevelI+T→T
Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language ModelsEMNLP 2024 Findings2024/07/31GithubEncoder LevelI+T→T
Sim-clip: Unsupervised siamese adversarial fine-tuning for robust and semantically-rich vision-language modelsArxiv 20242024/07/20GithubEncoder LevelI+T→T
Can Textual Unlearning Solve Cross-Modality Safety Alignment?EMNLP 2024 Findings2024/05/27NoneGenerator LevelI+T→T
Safety alignment for vision language modelsArxiv 20242024/05/22NoneGenerator LevelI+T→T
Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield promptingECCV 20242024/05/14GithubInput LevelI+T→T
Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text TransformationECCV 20242024/03/14GithubOutput LevelI+T→T
Safety fine-tuning at (almost) no cost: A baseline for vision large language modelsICML 20242024/02/03GithubGenerator LevelI+T→T
Inferaligner: Inference-time alignment for harmlessness through cross-model guidanceEMNLP 20242024/01/20GithubGenerator LevelI+T→T
Mllm-protector: Ensuring mllm’s safety without hurting performanceEMNLP 20242024/01/05GithubOutput LevelI+T→T
Jailguard: A universal detection framework for llm prompt-based attacksArxiv 20232023/12/17GithubOutput LevelI+T→T
Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructionsICLR 20242023/09/14GithubGenerator LevelI+T→T

Jailbreak Defense of Any-to-Vision Models

TitleVenueDateCodeTaxonomyMultimodal Model
Safe-Control: A Safety Patch for Mitigating Unsafe Content in Text-to-Image Generation ModelsArxiv 20252025/08/28None---T→I
Seeing It Before It Happens: In-Generation NSFW Detection for Diffusion-Based Text-to-Image ModelsArxiv 20252025/08/05None---T→I
PromptSafe: Gated Prompt Tuning for Safe Text-to-Image GenerationArxiv 20252025/08/02None---T→I
Wukong Framework for Not Safe For Work Detection in Text-to-Image systemsArxiv 20252025/08/01None---T→I
NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image GenerationArxiv 20252025/07/23None---T→I
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video ModelsArxiv 20252025/04/22None---T→V
Towards NSFW-Free Text-to-Image Generation via Safety-Constraint Direct Preference OptimizationArxiv 20252025/04/19None---T→I
I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video ModelsCVPR 2025---None---T→V
Hyperbolic Safety-Aware Vision-Language ModelsCVPR 20252025/03/15Github---T→I
Distorting Embedding Space for Safety: A Defense Mechanism for Adversarially Robust Diffusion ModelsArxiv 20252025/03/10Github---T→I
SafeText: Safe Text-to-image Models via Aligning the Text EncoderArxiv 20252025/02/28None---T→I
Comprehensive Assessment and Analysis for NSFW Content Erasure in Text-to-Image Diffusion ModelsArxiv 20252025/02/18None---T→I
A Comprehensive Survey on Concept Erasure in Text-to-Image Diffusion ModelsArxiv 20252025/02/17None---T→I
Training-Free Safe Denoisers for Safe Use of Diffusion ModelsArxiv 20252025/02/11None---T→I
Beautiful Images, Toxic Words: Understanding and Addressing Offensive Text in Generated ImagesArxiv 20252025/02/07None---T→I
Distorting Embedding Space for Safety: A Defense Mechanism for Adversarially Robust Diffusion ModelsArxiv 20252025/01/30Github---T→I
CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word VocabularyArxiv 20252025/01/26None---T→I
CROPS: Model-Agnostic Training-Free Framework for Safe Image Synthesis with Latent Diffusion ModelsArxiv 20252025/01/09None---T→I
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image ModelsArxiv 20252025/01/07Homepage---T→I
DuMo: Dual Encoder Modulation Network for Precise Concept ErasureAAAI 20252025/01/02Github---T→I
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image ModelsArxiv 20242024/12/24None---T→I
SafeCFG: Redirecting Harmful Classifier-Free Guidance for Safe GenerationArxiv 20242024/12/20None---T→I
SafetyDPO: Scalable Safety Alignment for Text-to-Image GenerationICCV 20252024/12/13Github---T→I
TraSCE: Trajectory Steering for Concept ErasureArxiv 20242024/12/10Github---T→I
Buster: Incorporating Backdoor Attacks into Text Encoder to Mitigate NSFW Content GenerationArxiv 20242024/12/10None---T→I
Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise OptimizationArxiv 20242024/12/05Github---T→I
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion ModelsArxiv 20242024/11/30None---T→I
Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent ReconstructionArxiv 20242024/11/21None---T→I
Safe Text-to-Image Generation:Simply Sanitize the Prompt EmbeddingArxiv 20242024/11/15NoneEncoder LevelT→I
Safree: Training-free and adaptive guard for safe text-to-image and video generationICLR 20252024/10/16GithubGenerator LevelT→I/T→V
Shielddiff: Suppressing sexual content generation from diffusion models through reinforcement learningArxiv 20242024/10/04GithubGenerator LevelT→I
Dark miner: Defend against unsafe generation for text-to-image diffusion modelsArxiv 20242024/09/26NoneGenerator LevelT→I
Score forgetting distillation: A swift, data-free method for machine unlearning in diffusion modelsICLR 20252024/09/17NoneGenerator LevelT→I
EIUP: A Training-Free Approach to Erase Non-Compliant Concepts Conditioned on Implicit Unsafe PromptsArxiv 20242024/08/02NoneGenerator LevelT→I
Direct Unlearning Optimization for Robust and Safe Text-to-Image ModelsNeurIPS 20242024/07/17GithubGenerator LevelT→I
Reliable and Efficient Concept Erasure of Text-to-Image Diffusion ModelsECCV 20242024/07/17GithubGenerator LevelT→I
Conceptprune: Concept editing in diffusion models via skilled neuron pruningArxiv 20242024/05/29GithubGenerator LevelT→I
Pruning for Robust Concept Erasing in Diffusion ModelsNeurIPS SafeGenAi Workshop 20242024/05/26NoneGenerator LevelT→I
Defensive unlearning with adversarial training for robust concept erasure in diffusion modelsNeurIPS 20242024/05/24GithubEncoder LevelT→I
Unlearning concepts in diffusion model via concept domain correction and concept preserving gradientAAAI 20252024/05/24GithubGenerator LevelT→I
Espresso: Robust Concept Filtering in Text-to-Image ModelsCODASPY 20252024/04/30NoneOutput LevelT→I
Latent Guard: a Safety Framework for Text-to-image GenerationECCV 20242024/04/11GithubEncoder LevelT→I
SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image ModelsACM CCS 20242024/04/10GithubGenerator LevelT→I
Salun: Empowering machine unlearning via gradient-based weight saliency in both image classification and generationICLR 20242024/04/04GithubGenerator LevelT→I
GuardT→I: Defending Text-to-Image Models from Adversarial PromptsNeurIPS 20242024/03/03NoneEncoder LevelT→I
Universal prompt optimizer for safe text-to-image generationNAACL 20242024/02/16GithubInput LevelT→I
Erasediff: Erasing data influence in diffusion modelsArxiv 20242024/01/11GithubGenerator LevelT→I
Localization and manipulation of immoral visual cues for safe text-to-image generationWACV 20242024/01/01NoneOutput LevelT→I
Receler: Reliable concept erasing of text-to-image diffusion models via lightweight erasersECCV 20242023/11/29GithubGenerator LevelT→I
Self-discovering interpretable diffusion latent directions for responsible text-to-image generationCVPR 20242023/11/28GithubEncoder LevelT→I
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language ModelsECCV 20242023/11/27GithubEncoder LevelT→I
Mace: Mass concept erasure in diffusion modelsCVPR 20242023/10/19GithubGenerator LevelT→I
Implicit concept removal of diffusion modelsECCV 20242023/10/09NoneInput LevelT→I
Unified concept editing in diffusion modelsWACV 20242023/08/25GithubGenerator LevelT→I
Towards safe self-distillation of internet-scale text-to-image diffusion modelsICML 2023 Workshop on Challenges in Deployable Generative AI2023/07/12GithubGenerator LevelT→I
Forget-Me-Not: Learning to Forget in Text-to-Image Diffusion ModelsCVPR Workshop 20242023/05/30GithubGenerator LevelT→I
Erasing concepts from diffusion modelsICCV 20232023/05/13GithubGenerator LevelT→I
Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion ModelsCVPR 20232022/11/09GithubGenerator LevelT→I

Jailbreak Defense of Any-to-Any Models

TitleVenueDateCodeTaxonomyMultimodal Model

💯Evaluation

⭐️Evaluation Datasets

Below is a comparison table of publicly available representative evaluation datasets and a description of each attribute in the table.

  • Collected: raw data created by humans or collected from real-world websites.
  • Reconstructed: Data reorganized from other existing datasets.
  • Synthesized: AI-generated data using LLM or diffusion models.
  • Adversarial: Adversarial data generated by jailbreak attack methods.

Used to Any-to-Text Models

DatasetText SourceImage SourceVolumeThemeAccess
FigstepSynthesizedAdversarial50010Github
AdvBenchSynthesized---500---Github
ReadTeam-2KCollected & Reconstructed & SynthesizedN/A200016Huggingface
HarmBenchCollected---5104Github
HADESSynthesizedCollected & Synthesized & Adversarial7505Github
MM-SafetyBenchSynthesizedSynthesized & Adversarial504013Github
JailBreakV-28KAdversarialReconstructed & Synthesized2800016Huggingface

Used to Any-to-Vision Models

DatasetText SourceImage SourceVolumeAccessTheme
NSFW-200Synthesized---200---Github
MMAReconstructed & AdversarialAdversarial1000---Huggingface
VBCDEReconstructed & Adversarial---1005Github
I2PCollectedCollected47037Huggingface
Unsafe DiffusionCollected & Reconstructed---1434---Github
MACE-CelebrityCollected---1000---Github
MACE-ArtReconstructed---1000---Github
MPUPSynthesized---12004Huggingface
T2VSafetyBenchReconstructed & Synthesized & Adversarial---440012Github

📚Evaluation Methods

Current evaluation methods are primarily classified into two categories: manual evaluation and automated evaluation.

  • Manual evaluation involves human assessment to determine if the content is toxic, offering a direct and interpretable method of evaluation.
  • Automated approaches assess the safety of multimodal generative models by employing a range of techniques, including detector-based, GPT-based, and rule-based methods.
jailbreak_evaluation

Text Detector

Toxicity detectorAccess
LLama-GuardHuggingface
LLama-Guard2Huggingface
DetoxifyGithub
GPTFUZZERHuggingface
Perspective APIWebsite

Image Detector

Toxicity detectorAccess
NudeNetGithub
Q16Github
Safety CheckerHuggingface
ImgcensorGithub
Multi-headed Safety ClassifierGithub

😉Citation

If you find this work useful in your research, Please kindly cite using the following BibTex:

@article{liu2024jailbreak,
    title={Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey},
    author={Liu, Xuannan and Cui, Xing and Li, Peipei and Li, Zekun and Huang, Huaibo and Xia, Shuhan and Zhang, Miaoxuan and Zou, Yueying and He, Ran},
    journal={arXiv preprint arXiv:2411.09259},
    year={2024},
}

Contributors

shuhanxia

357 commits

liuxuannan

36 commits

NST666

7 commits

liuxuannan/Awesome-Multimodal-Jailbreak

A Survey on Jailbreak Attacks and Defenses against Multimodal Generative Models

335

403 commits

updated Jan 11, 2026

See the code

README

😈🛡️Awesome-Jailbreak-against-Multimodal-Generative-Models

🔥🔥🔥 Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey

Paper

We've curated a collection of the latest 😋, most comprehensive 😎, and most valuable 🤩 resources on Jailbreak Attack and Defense against Multimodel Generative Models.
But we don't stop there; Our repository is constantly updated to ensure you have the most current information at your fingertips.

survey model

🤗Introduction

This survey presents a comprehensive review of existing jailbreak attack and defense against multimodal generative models.
Given the generalized lifecycle of multimodal jailbreak, we systematically explore attacks and corresponding defense strategies across four levels: input, encoder, generator, and output.

🧑‍💻 Four Levels of Multimodal Jailbreak lifecycle

  • Input Level: Attackers and defenders operate solely on the input data. Attackers modify inputs to execute attacks, while defenders incorporate protective cues to enhance detection.
  • Encoder Level: With access to the encoder, attackers optimize adversarial inputs to inject malicious information into the encoding process, while defenders work to prevent harmful information from being encoded within the latent space.
  • Generator Level: : With full access to the generative models, attackers leverage inference information, such as activations and gradients, and fine-tune models to increase adversarial effectiveness, while defenders use these techniques to strengthen model robustness.
  • Output Level: With the output from the generative model, attackers can iteratively refine adversarial inputs, while defenders can apply post-processing techniques to enhance detection.

Based on this analysis, we present a detailed taxonomy of attack methods, defense mechanisms, and evaluation frameworks specific to multimodal generative models.
We cover a wide range of input-output configurations, including modalities such as Any-to-Text, Any-to-Vision, and Any-to-Any within generative systems.

survey model

🚀Table of Contents

🔥Multimodal Generative Models

Below are tables of model short name and representative generative models used for jailbreak. For input/output modalities, I: Image, T: Text, V: Video, A: Audio.

📑Any-to-Text Models (LLM Backbone)

Short NameModalityRepresentative Model
I+T→TI + T → TLLaVA, MiniGPT4, InstructBLIP
VT2TV + T → TVideo-LLaVA, Video-LLaMA
AT2TA + T → TAudio Flamingo, Audiopalm

📖Any-to-Vision (Diffusion Backbone)

Short NameModalityRepresentative Model
T→IT → IStable Diffusion, Midjourney, DALLE
IT→II + T → IDreamBooth, InstructP2P
T2VT → VOpen-Sora, Stable Video Diffusion
IT2VI + T → VVideoPoet, CogVideoX

📰Any-to-Any (Unified Backbone)

Short NameModalityRepresentative Model
IT→ITI + T → I + TNext-GPT, Chameleon
TIV2TIVT + I + V → T + I + VEMU3
Any2AnyAny → AnyGPT-4o, Gemini Ultra

😈JailBreak Attack

📖Attack-Intro

We categorize attack methods into black-box, gray-box, and white-box attacks. in a black-box setting where the model is inaccessible to the attacker, the attack is limited to surface-level interactions, focusing solely on the model’s input and/or output. Regarding gray-box and white-box attacks, we consider model-level attacks, including attacks at both the encoder and generator.

  • Input-level attack: attackers are compelled to develop more sophisticated input templates across prompt engineering, image engineering, and role-ploy techniques.
  • Output-level attack: Attackers focus on querying outputs across multiple input variants. Driven by specific adversarial goals, attackers employ estimation-based and search-based attack techniques to iteratively refine these input variants.
jailbreak_attack_black_box
  • Encoder-level attack: Attackers are restricted to accessing only the encoders to provoke harmful responses. In this case, attackers typically seek to maximize cosine similarity within the latent space, ensuring the adversarial input retains similar semantics to the target malicious content while still being classified as safe.
  • Generator-level attack: Attackers have unrestricted access to the generative model’s architecture and checkpoint, enabling attackers to conduct thorough investigations and manipulations, thus enabling sophisticated attacks.
jailbreak_attack_white_and_gray_box

📑Papers

Below are the papers related to jailbreak attacks.

Jailbreak Attack of Any-to-Text Models

TitleVenueDateCodeTaxonomyMultimodal Model
Jailbreaking Large Vision Language Models in Intelligent Transportation SystemsArxiv 20252025/11/17None---I+T→T
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMsArxiv 20252025/11/20None---I+T→T
The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental JailbreaksArxiv 20252025/11/20None---I+T→T
Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMsNeurIPS 20252025/05/17HomepageInput LevelV+T→T
Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-IntensityArxiv 20252025/08/11Github---I+T→T
JPS: Jailbreak Multimodal Large Language Models with Collaborative Visual Perturbation and Textual SteeringACM MM 20252025/08/07Github---I+T→T
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM JailbreakingArxiv 20252025/07/29None---I+T→T
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language ModelsArxiv 20252025/07/20None---I+T→T
Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language ModelsTMM 20252025/07/18None---I+T→T
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context InjectionArxiv 20252025/07/03Github---I+T→T
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language ModelsTCSVT 20252025/07/02Github---I+T→T
USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language ModelsArxiv 20252025/05/26Github---I+T→T
VSCBench: Bridging the Gap in Vision-Language Model Safety CalibrationArxiv 20252025/05/26Github---I+T→T
JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language ModelsArxiv 20252025/05/26None---I+T→T
Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language ModelsArxiv 20252025/05/26None---A+T→T
Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box FrameworArxiv 20252025/05/24None---A+T→T
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language ModelsArxiv 20252025/05/23None---A+T→T
BadNAVer: Exploring Jailbreak Attacks On Vision-and-Language NavigationArxiv 20252025/05/22None---I+T→T
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language ModelsArxiv 20252025/05/20None---A+T→T
Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language ModelsArxiv 20252025/05/18None---I+T→T
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning ModelArxiv 20252025/05/10Github---I+T→T
SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning ModelsArxiv 20252025/04/09Github---I+T→T
PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code ContextualizationArxiv 20252025/04/02None---I+T→T
Multilingual and Multi-Accent Jailbreaking of Audio LLMsArxiv 20252025/04/01None---A+T→T
Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution StrategyCVPR 20252025/03/26Github---I+T→T
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak AttacksArxiv 20252025/03/24None---I+T→T
Making Every Step Effective: Jailbreaking Large Vision-Language Models Through Hierarchical KV EqualizationArxiv 20252025/03/14None---I+T→T
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist ContentArxiv 20252025/03/13None---I+T→T
Utilizing Jailbreak Probability to Attack and Safeguard Multimodal LLMsArxiv 20252025/03/10None---I+T→T
FC-Attack: Jailbreaking Large Vision-Language Models via Auto-Generated FlowchartsArxiv 20252025/02/28None---I+T→T
EigenShield: Causal Subspace Filtering via Random Matrix Theory for Adversarially Robust Vision-Language ModelsArxiv 20252025/02/20None---I+T→T
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMsArxiv 20252025/02/16None---I+T→T
Distraction is All You Need for Multimodal Large Language Model JailbreakingCVPR 20252025/02/15None---I+T→T
ELITE: Enhanced Language-Image Toxicity Evaluation for SafetyArxiv 20252025/02/07None---I+T→T
`Do as I say not as I do': A Semi-Automated Approach for Jailbreak Prompt Attack against Multimodal LLMsArxiv 20252025/02/02None---A+T→T
"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language ModelsArxiv 20252025/02/02Github---A+T→T
Failures to Find Transferable Image Jailbreaks Between Vision-Language ModelsICLR 20252025/01/23GithubGenerator LevelI+T→T
Jailbreaking Multimodal Large Language Models via Shuffle InconsistencyICCV 20252025/01/09None---I+T→T
Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language ModelsArxiv 20242024/12/21None---I+T+A→T
AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language ModelsICLR 20252024/12/11GithubGenerator LevelA+T→T
Heuristic-Induced Multimodal Risk Distribution Jailbreak Attack for Multimodal Large Language ModelsICCV 20252024/12/8Github---I+T→T
PBI-Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity MaximizationArxiv 20242024/12/8None---I+T→T
Jailbreak Large Vision-Language Models Through Multi-Modal LinkageArxiv 20242024/11/30Github---I+T→T
Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language ModelsCVPR 20252024/11/27None---I+T→T
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and DefenseArxiv 20242024/11/13None---I+T→T
MMJ-Bench : A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language ModelsArxiv 20242024/08/16None---I+T→T
Failures to Find Transferable Image Jailbreaks Between Vision-Language ModelsNeurIPS 2024 Workshops2024/07/21None---I+T→T
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language ModelsNeurIPS 20242024/06/11Github---I+T→T
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak AttacksArxiv 20242024/06/10Github---I+T→T
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?Arxiv 20242024/04/04Github---I+T→T
VLSBench: Unveiling Visual Leakage in Multimodal SafetyACL 20252024/11/29HomepageInput LevelI+T→T
Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language ModelsArxiv 20242024/11/18GithubOutput LevelI+T→T
IDEATOR: Jailbreaking Large Vision-Language Models Using ThemselvesICCV 20252024/11/15GithubOutput LevelI+T→T
Zer0-Jack: A memory-efficient gradient-based jailbreaking method for black box Multi-modal Large Language ModelsNeurIPS SafeGenAi Workshop 20242024/11/12NoneOutput LevelI+T→T
Audio is the achilles’heel: Red teaming audio large multimodal modelsArxiv 20242024/10/31GithubInput LevelA+T→T
Advweb: Controllable black-box attacks on vlm-powered web agentsArxiv 20242024/10/22NoneInput LevelI+T→T
Can Large Language Models Automatically Jailbreak GPT-4V?NAACL Workshop 20242024/07/23NoneInput LevelI+T→T
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak PromptsACM MM 20242024/07/21NoneInput LevelI+T→T
Image-to-Text Logic Jailbreak: Your Imagination can Help You Do AnythingArxiv 20242024/07/01NoneInput LevelI+T→T
From LLMs to MLLMs: Exploring the Landscape of Multimodal JailbreakingEMNLP 20242024/06/21NoneEncoder LevelI+T→T
Jailbreak Vision Language Models via Bi-Modal Adversarial PromptArxiv 20242024/06/06GithubGenerator LevelI+T→T
Efficient LLM-Jailbreaking by Introducing Visual ModalityArxiv 20242024/05/30GithubGenerator LevelI+T→T
White-box Multimodal Jailbreaks Against Large Vision-Language ModelsACM MM 20242024/05/28GithubGenerator LevelI+T→T
Medical MLLM is Vulnerable: Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language ModelsArxiv 20242024/05/26Github---I+T→T
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image CharacterArxiv 20242024/05/25GithubInput LevelI+T→T
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language ModelsECCV 20242024/05/14GithubGenerator LevelI+T→T
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially FastICML 20242024/02/13GithubGenerator LevelI+T→T
Jailbreaking Attack against Multimodal Large Language ModelArxiv 20242024/02/04GithubGenerator LevelI+T→T
Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsICLR 2024 Spotlight2024/01/16GithubEncoder LevelI+T→T
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language ModelsECCV 20242023/11/29GithubInput LevelI+T→T
How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMsECCV 20242023/11/27GithubEncoder LevelI+T→T
Jailbreaking GPT-4V via Self-Adversarial Attacks with System PromptsArxiv 20232023/11/15NoneOutput LevelI+T→T
FigStep: Jailbreaking Large Vision-language Models via Typographic Visual PromptsAAAI 20252023/11/09GithubInput LevelI+T→T
Image Hijacks: Adversarial Images can Control Generative Models at RuntimeICML 20242023/09/01GithubGenerator LevelI+T→T
Are aligned neural networks adversarially aligned?NeurIPS 20232023/06/26NoneGenerator LevelI+T→T
Visual Adversarial Examples Jailbreak Aligned Large Language ModelsAAAI 20242023/06/22GithubGenerator LevelI+T→T
On Evaluating Adversarial Robustness of Large Vision-Language ModelsNeurIPS 20232023/05/26HomepageEncoder LevelI+T→T

Jailbreak Attack of Any-to-Vision Models

TitleVenueDateCodeTaxonomyMultimodal Model
Universally Unfiltered and Unseen:Input-Agnostic Multimodal Jailbreaks against Text-to-Image Model SafeguardsACM MM 20252025/07/30Github---T→I
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image ModelsArxiv 20252025/07/23None---T→I
PLA: Prompt Learning Attack against Text-to-Image Generative ModelsICCV 20252025/07/14None---T→I
GhostPrompt: Jailbreaking Text-to-image Generative Models based on Dynamic OptimizationArxiv 20252025/05/25None---T→I
TokenProber: Jailbreaking Text-to-image Models via Fine-grained Word Impact AnalysisArxiv 20252025/05/11None---T→I
T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak AttacksArxiv 20252025/05/10None---T→V
Inception: Jailbreak the Memory Mechanism of Text-to-Image Generation SystemsArxiv 20252025/04/29None---T→I
Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image ModelsArxiv 20252025/04/15None---T→I
Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive JailbreakingCVPR 2025 Highlight2025/04/08Github---T→I
Reason2Attack: Jailbreaking Text-to-Image Models via LLM ReasoningArxiv 20252025/03/23None---T→I
Jailbreaking Safeguarded Text-to-Image Models via Large Language ModelsArxiv 20252025/03/03None---T→I
Unified Prompt Attack Against Text-to-Image Generation ModelsTPAMI 20252024/02/23None---T→I
T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image GenerationArxiv 20252025/02/22Github---T→I
CogMorph: Cognitive Morphing Attacks for Text-to-Image ModelsArxiv 20252024/01/21None---T→I
FameBias: Embedding Manipulation Bias Attack in Text-to-Image ModelsArxiv 20242024/12/24None---T→I
Antelope: Potent and Concealed Jailbreak Attack StrategyArxiv 20242024/12/11None---T→I
Multimodal Pragmatic Jailbreak on Text-to-image ModelsArxiv 20242024/09/27None---T→I
In-Context Experience Replay Facilitates Safety Red-Teaming of Text-to-Image Diffusion ModelsArxiv 20242024/11/25NoneOutput LevelT→I
Unfiltered and Unseen: Universal Multimodal Jailbreak Attacks on Text-to-Image Model DefensesOpenreview2024/11/13None---T→I
AdvI2I: Adversarial Image Attack on Image-to-Image Diffusion modelsArxiv 20242024/10/28GithubEncoder LevelT→I
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by StepArxiv 20242024/10/4NoneOutput LevelT→I
ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep GenerationNeurIPS 20242024/9/25GithubInput LevelT→I
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image ModelsArxiv 20242024/08/25NoneOutput LevelT→I
Perception-guided Jailbreak against Text-to-Image ModelsAAAI 20252024/08/20GithubInput LevelT→I
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order OptimizationNAACL 20252024/08/18GithubOutput LevelT→I
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion ModelsArxiv 20242024/08/02NoneEncoder LevelT→I
Jailbreaking Text-to-Image Models with LLM-Based AgentsArxiv 20242024/08/01NoneOutput LevelT→I
Automatic Jailbreaking of the Text-to-Image Generative AI SystemsICML 2024 Workshop NextGenAISafety2024/05/26GithubOutput LevelT→I
UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual CheckersICML 20242024/05/18NoneInput LevelT→I
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image GeneratorsArxiv 20242024/02/23NoneInput LevelT→I
Harnessing LLM to Attack LLM-Guarded Text-to-Image ModelsArxiv 20232023/12/12GithubInput LevelT→I
MMA-Diffusion: MultiModal Attack on Diffusion ModelsCVPR 20242023/11/29GithubEncoder LevelT→I
VA3: Virtually Assured Amplification Attack on Probabilistic Copyright Protection for Text-to-Image Generative ModelsCVPR 20242023/11/29GithubGenerator LevelT→I
To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For NowECCV 20242023/10/18GithubGenerator LevelT→I
Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?ICLR 20242023/10/16GithubEncoder LevelT→I
SurrogatePrompt: Bypassing the Safety Filter of Text-To-Image Models via SubstitutionCCS 20242023/09/25NoneInput LevelT→I
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic PromptsICML 20242023/09/12GithubGenerator LevelT→I
SneakyPrompt: Jailbreaking Text-to-image Generative ModelsSymposium on Security and Privacy 20242023/05/20GithubOutput LevelT→I
Red-Teaming the Stable Diffusion Safety FilterNeurIPSW 20222022/10/03NoneInput LevelT→I

Jailbreak Attack of Any-to-Any Models

TitleVenueDateCodeTaxonomyMultimodal Model
Gradient-based Jailbreak Images for Multimodal Fusion ModelsArxiv 20242024/10/4GithubGenerator LevelI+T→I+T
Voice jailbreak attacks against gpt-4oArxiv 20242024/05/29GithubOutput LevelAny→Any

🛡️Jailbreak Defense

📖Defense-Intro

Current efforts made in the jailbreak defense of multimodal generative models include two lines of work: Discriminative defense and Transformative defense.

  • Discriminative defenses: is constrained to classification tasks for assigning binary labels.
jailbreak_discriminative_defense
  • Transformative Defense: aims to produce appropriate and safe responses in the presence of malicious or adversarial inputs.
jailbreak_transformative_defense

📑Papers

Below are the papers related to jailbreak defense.

Jailbreak Defense of Any-to-Text Models

TitleVenueDateCodeTaxonomyMultimodal Model
Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model SecurityArxiv 20252025/11/20Github---I+T→T
Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models: A Unified and Accurate ApproachArxiv 20252025/08/08None---I+T→T
Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model SecurityArxiv 20252025/07/29None---I+T→T
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore MechanismArxiv 20252025/07/02None---I+T→T
The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language ModelsArxiv 20252025/06/15None---I+T→T
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language ModelsArxiv 20252025/05/28None---I+T→T
GuardReasoner-VL: Safeguarding VLMs via Reinforced ReasoningArxiv 20252025/05/16Github---I+T→T
DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language ModelsArxiv 20252025/04/25Github---I+T→T
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?CVPR 20252025/04/14None---I+T→T
JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language ModelArxiv 20252025/04/3Github---I+T→T
Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based AttacksICCV 20252025/04/2Github---I+T→T
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial DefenseArxiv 20252025/03/14None---I+T→T
Utilizing Jailbreak Probability to Attack and Safeguard Multimodal LLMsArxiv 20252025/03/10None---I+T→T
Adversarial Training for Multimodal Large Language Models against Jailbreak AttacksArxiv 20252025/03/05None---I+T→T
HiddenDetect: Detecting Jailbreak Attacks against Large Vision-Language Models via Monitoring Hidden StatesACL 20252025/02/20Github---I+T→T
SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine UnlearningArxiv 20252025/02/18None---I+T→T
Understanding and Rectifying Safety Perception Distortion in VLMsArxiv 20252025/02/18None---I+T→T
Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial TrainingArxiv 20252025/02/17None---I+T→T
Towards Robust Multimodal Large Language Models Against Jailbreak AttacksArxiv 20252025/02/02None---I+T→T
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language ModelsArxiv 20252025/01/30Github---I+T→T
Internal Activation Revision: Safeguarding Vision Language Models Without Parameter UpdateArxiv 20252025/01/24None---I+T→T
MSTS: A Multimodal Safety Test Suite for Vision-Language ModelsArxiv 20252025/01/17Github---I+T→T
Spot Risks Before Speaking! Unraveling Safety Attention Heads in Large Vision-Language ModelsArxiv 20252025/01/03Github---I+T→T
Defending LVLMs Against Vision Attacks through Partial-Perception SupervisionArxiv 20242024/12/17None---I+T→T
VLMGuard: Defending VLMs against Malicious Prompts via Unlabeled DataArxiv 20242024/10/01None---I+T→T
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time AlignmentCVPR 20252024/11/27GithubOutput LevelI+T→T
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against JailbreaksCVPR 20252024/11/23GithubGenerator LevelI+T→T
Uniguard: Towards universal safety guardrails for jailbreak attacks on multimodal large language modelsArxiv 20242024/11/03NoneInput LevelI+T→T
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single VectorArxiv 20242024/10/30GithubGenerator LevelI+T→T
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak AttacksICLR 20252024/10/28GithubInput LevelI+T→T
The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?Arxiv 20242024/10/02NoneInput LevelI+T→T
CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional CalibrationCOLM 20242024/9/17NoneOutput LevelI+T→T
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial AttacksBigData 20242024/09/11NoneEncoder LevelI+T→T
Bathe: Defense against the jailbreak attack in multimodal large language models by treating harmful instruction as backdoor triggerArxiv 20242024/08/17NoneGenerator LevelI+T→T
Cross-modality Information Check for Detecting Jailbreaking in Multimodal Large Language ModelsEMNLP 2024 Findings2024/07/31GithubEncoder LevelI+T→T
Sim-clip: Unsupervised siamese adversarial fine-tuning for robust and semantically-rich vision-language modelsArxiv 20242024/07/20GithubEncoder LevelI+T→T
Can Textual Unlearning Solve Cross-Modality Safety Alignment?EMNLP 2024 Findings2024/05/27NoneGenerator LevelI+T→T
Safety alignment for vision language modelsArxiv 20242024/05/22NoneGenerator LevelI+T→T
Adashield: Safeguarding multimodal large language models from structure-based attack via adaptive shield promptingECCV 20242024/05/14GithubInput LevelI+T→T
Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text TransformationECCV 20242024/03/14GithubOutput LevelI+T→T
Safety fine-tuning at (almost) no cost: A baseline for vision large language modelsICML 20242024/02/03GithubGenerator LevelI+T→T
Inferaligner: Inference-time alignment for harmlessness through cross-model guidanceEMNLP 20242024/01/20GithubGenerator LevelI+T→T
Mllm-protector: Ensuring mllm’s safety without hurting performanceEMNLP 20242024/01/05GithubOutput LevelI+T→T
Jailguard: A universal detection framework for llm prompt-based attacksArxiv 20232023/12/17GithubOutput LevelI+T→T
Safety-tuned llamas: Lessons from improving the safety of large language models that follow instructionsICLR 20242023/09/14GithubGenerator LevelI+T→T

Jailbreak Defense of Any-to-Vision Models

TitleVenueDateCodeTaxonomyMultimodal Model
Safe-Control: A Safety Patch for Mitigating Unsafe Content in Text-to-Image Generation ModelsArxiv 20252025/08/28None---T→I
Seeing It Before It Happens: In-Generation NSFW Detection for Diffusion-Based Text-to-Image ModelsArxiv 20252025/08/05None---T→I
PromptSafe: Gated Prompt Tuning for Safe Text-to-Image GenerationArxiv 20252025/08/02None---T→I
Wukong Framework for Not Safe For Work Detection in Text-to-Image systemsArxiv 20252025/08/01None---T→I
NSFW-Classifier Guided Prompt Sanitization for Safe Text-to-Image GenerationArxiv 20252025/07/23None---T→I
T2VShield: Model-Agnostic Jailbreak Defense for Text-to-Video ModelsArxiv 20252025/04/22None---T→V
Towards NSFW-Free Text-to-Image Generation via Safety-Constraint Direct Preference OptimizationArxiv 20252025/04/19None---T→I
I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video ModelsCVPR 2025---None---T→V
Hyperbolic Safety-Aware Vision-Language ModelsCVPR 20252025/03/15Github---T→I
Distorting Embedding Space for Safety: A Defense Mechanism for Adversarially Robust Diffusion ModelsArxiv 20252025/03/10Github---T→I
SafeText: Safe Text-to-image Models via Aligning the Text EncoderArxiv 20252025/02/28None---T→I
Comprehensive Assessment and Analysis for NSFW Content Erasure in Text-to-Image Diffusion ModelsArxiv 20252025/02/18None---T→I
A Comprehensive Survey on Concept Erasure in Text-to-Image Diffusion ModelsArxiv 20252025/02/17None---T→I
Training-Free Safe Denoisers for Safe Use of Diffusion ModelsArxiv 20252025/02/11None---T→I
Beautiful Images, Toxic Words: Understanding and Addressing Offensive Text in Generated ImagesArxiv 20252025/02/07None---T→I
Distorting Embedding Space for Safety: A Defense Mechanism for Adversarially Robust Diffusion ModelsArxiv 20252025/01/30Github---T→I
CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word VocabularyArxiv 20252025/01/26None---T→I
CROPS: Model-Agnostic Training-Free Framework for Safe Image Synthesis with Latent Diffusion ModelsArxiv 20252025/01/09None---T→I
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image ModelsArxiv 20252025/01/07Homepage---T→I
DuMo: Dual Encoder Modulation Network for Precise Concept ErasureAAAI 20252025/01/02Github---T→I
AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image ModelsArxiv 20242024/12/24None---T→I
SafeCFG: Redirecting Harmful Classifier-Free Guidance for Safe GenerationArxiv 20242024/12/20None---T→I
SafetyDPO: Scalable Safety Alignment for Text-to-Image GenerationICCV 20252024/12/13Github---T→I
TraSCE: Trajectory Steering for Concept ErasureArxiv 20242024/12/10Github---T→I
Buster: Incorporating Backdoor Attacks into Text Encoder to Mitigate NSFW Content GenerationArxiv 20242024/12/10None---T→I
Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise OptimizationArxiv 20242024/12/05Github---T→I
Safety Alignment Backfires: Preventing the Re-emergence of Suppressed Concepts in Fine-tuned Text-to-Image Diffusion ModelsArxiv 20242024/11/30None---T→I
Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent ReconstructionArxiv 20242024/11/21None---T→I
Safe Text-to-Image Generation:Simply Sanitize the Prompt EmbeddingArxiv 20242024/11/15NoneEncoder LevelT→I
Safree: Training-free and adaptive guard for safe text-to-image and video generationICLR 20252024/10/16GithubGenerator LevelT→I/T→V
Shielddiff: Suppressing sexual content generation from diffusion models through reinforcement learningArxiv 20242024/10/04GithubGenerator LevelT→I
Dark miner: Defend against unsafe generation for text-to-image diffusion modelsArxiv 20242024/09/26NoneGenerator LevelT→I
Score forgetting distillation: A swift, data-free method for machine unlearning in diffusion modelsICLR 20252024/09/17NoneGenerator LevelT→I
EIUP: A Training-Free Approach to Erase Non-Compliant Concepts Conditioned on Implicit Unsafe PromptsArxiv 20242024/08/02NoneGenerator LevelT→I
Direct Unlearning Optimization for Robust and Safe Text-to-Image ModelsNeurIPS 20242024/07/17GithubGenerator LevelT→I
Reliable and Efficient Concept Erasure of Text-to-Image Diffusion ModelsECCV 20242024/07/17GithubGenerator LevelT→I
Conceptprune: Concept editing in diffusion models via skilled neuron pruningArxiv 20242024/05/29GithubGenerator LevelT→I
Pruning for Robust Concept Erasing in Diffusion ModelsNeurIPS SafeGenAi Workshop 20242024/05/26NoneGenerator LevelT→I
Defensive unlearning with adversarial training for robust concept erasure in diffusion modelsNeurIPS 20242024/05/24GithubEncoder LevelT→I
Unlearning concepts in diffusion model via concept domain correction and concept preserving gradientAAAI 20252024/05/24GithubGenerator LevelT→I
Espresso: Robust Concept Filtering in Text-to-Image ModelsCODASPY 20252024/04/30NoneOutput LevelT→I
Latent Guard: a Safety Framework for Text-to-image GenerationECCV 20242024/04/11GithubEncoder LevelT→I
SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image ModelsACM CCS 20242024/04/10GithubGenerator LevelT→I
Salun: Empowering machine unlearning via gradient-based weight saliency in both image classification and generationICLR 20242024/04/04GithubGenerator LevelT→I
GuardT→I: Defending Text-to-Image Models from Adversarial PromptsNeurIPS 20242024/03/03NoneEncoder LevelT→I
Universal prompt optimizer for safe text-to-image generationNAACL 20242024/02/16GithubInput LevelT→I
Erasediff: Erasing data influence in diffusion modelsArxiv 20242024/01/11GithubGenerator LevelT→I
Localization and manipulation of immoral visual cues for safe text-to-image generationWACV 20242024/01/01NoneOutput LevelT→I
Receler: Reliable concept erasing of text-to-image diffusion models via lightweight erasersECCV 20242023/11/29GithubGenerator LevelT→I
Self-discovering interpretable diffusion latent directions for responsible text-to-image generationCVPR 20242023/11/28GithubEncoder LevelT→I
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language ModelsECCV 20242023/11/27GithubEncoder LevelT→I
Mace: Mass concept erasure in diffusion modelsCVPR 20242023/10/19GithubGenerator LevelT→I
Implicit concept removal of diffusion modelsECCV 20242023/10/09NoneInput LevelT→I
Unified concept editing in diffusion modelsWACV 20242023/08/25GithubGenerator LevelT→I
Towards safe self-distillation of internet-scale text-to-image diffusion modelsICML 2023 Workshop on Challenges in Deployable Generative AI2023/07/12GithubGenerator LevelT→I
Forget-Me-Not: Learning to Forget in Text-to-Image Diffusion ModelsCVPR Workshop 20242023/05/30GithubGenerator LevelT→I
Erasing concepts from diffusion modelsICCV 20232023/05/13GithubGenerator LevelT→I
Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion ModelsCVPR 20232022/11/09GithubGenerator LevelT→I

Jailbreak Defense of Any-to-Any Models

TitleVenueDateCodeTaxonomyMultimodal Model

💯Evaluation

⭐️Evaluation Datasets

Below is a comparison table of publicly available representative evaluation datasets and a description of each attribute in the table.

  • Collected: raw data created by humans or collected from real-world websites.
  • Reconstructed: Data reorganized from other existing datasets.
  • Synthesized: AI-generated data using LLM or diffusion models.
  • Adversarial: Adversarial data generated by jailbreak attack methods.

Used to Any-to-Text Models

DatasetText SourceImage SourceVolumeThemeAccess
FigstepSynthesizedAdversarial50010Github
AdvBenchSynthesized---500---Github
ReadTeam-2KCollected & Reconstructed & SynthesizedN/A200016Huggingface
HarmBenchCollected---5104Github
HADESSynthesizedCollected & Synthesized & Adversarial7505Github
MM-SafetyBenchSynthesizedSynthesized & Adversarial504013Github
JailBreakV-28KAdversarialReconstructed & Synthesized2800016Huggingface

Used to Any-to-Vision Models

DatasetText SourceImage SourceVolumeAccessTheme
NSFW-200Synthesized---200---Github
MMAReconstructed & AdversarialAdversarial1000---Huggingface
VBCDEReconstructed & Adversarial---1005Github
I2PCollectedCollected47037Huggingface
Unsafe DiffusionCollected & Reconstructed---1434---Github
MACE-CelebrityCollected---1000---Github
MACE-ArtReconstructed---1000---Github
MPUPSynthesized---12004Huggingface
T2VSafetyBenchReconstructed & Synthesized & Adversarial---440012Github

📚Evaluation Methods

Current evaluation methods are primarily classified into two categories: manual evaluation and automated evaluation.

  • Manual evaluation involves human assessment to determine if the content is toxic, offering a direct and interpretable method of evaluation.
  • Automated approaches assess the safety of multimodal generative models by employing a range of techniques, including detector-based, GPT-based, and rule-based methods.
jailbreak_evaluation

Text Detector

Toxicity detectorAccess
LLama-GuardHuggingface
LLama-Guard2Huggingface
DetoxifyGithub
GPTFUZZERHuggingface
Perspective APIWebsite

Image Detector

Toxicity detectorAccess
NudeNetGithub
Q16Github
Safety CheckerHuggingface
ImgcensorGithub
Multi-headed Safety ClassifierGithub

😉Citation

If you find this work useful in your research, Please kindly cite using the following BibTex:

@article{liu2024jailbreak,
    title={Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey},
    author={Liu, Xuannan and Cui, Xing and Li, Peipei and Li, Zekun and Huang, Huaibo and Xia, Shuhan and Zhang, Miaoxuan and Zou, Yueying and He, Ran},
    journal={arXiv preprint arXiv:2411.09259},
    year={2024},
}

Contributors

shuhanxia

357 commits

liuxuannan

36 commits

NST666

7 commits