EddyLuo/JailBreakV_28K

Dataset

69

stars

114

commits

2

linked in READMEs

Jul 10, 2024

updated

README

⛓‍πŸ’₯ JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

🌐 GitHub | πŸ›Ž Project Page | πŸ‘‰ Download full datasets

If you like our project, please give us a star ⭐ on Hugging Face for the latest update.

πŸ“° News

DateEvent
2024/07/09πŸŽ‰ Our paper is accepted by COLM 2024.
2024/06/22πŸ› οΈ We have updated our version to V0.2, which supports users to customize their attack models and evaluate models.
2024/04/04🎁 We have posted our paper on Arxiv.
2024/04/03πŸŽ‰ We have released our evaluation and inference samples.
2024/03/30πŸ”₯ We have released our dataset.

πŸ“₯ Using our dataset via huggingface Dataset

from datasets import load_dataset


mini_JailBreakV_28K = load_dataset("JailbreakV-28K/JailBreakV-28k", 'JailBreakV_28K')["mini_JailBreakV_28K"]
JailBreakV_28K = load_dataset("JailbreakV-28K/JailBreakV-28k", 'JailBreakV_28K')["JailBreakV_28K"]
RedTeam_2K = load_dataset("JailbreakV-28K/JailBreakV-28k", 'RedTeam_2K')["RedTeam_2K"]

πŸ‘» Inference and Evaluation

Create environment

conda create -n jbv python=3.9
conda activate jbv
pip install -r requirements.txt

Conduct jailbreak attack on MLLMs

# we default use Bunny-v1_0, you can change the default attack model to your customized attack models by editing the annotated codes.
# You can follow the Bunny script in <attack_models> to add other attack models.
python attack.py --root JailBreakV_28K 

Conduct evaluation

# we default use LlamaGuard, you can change the default evaluate model to your customized evaluate models by editing the annotated codes.
# You can follow the LlamaGuard script in <evaluate_models> to add other evaluate models.
python eval.py --data_path ./results/JailBreakV_28k/<your customized attack model>/JailBreakV_28K.csv

πŸ˜ƒ Dataset Details

JailBreakV_28K and mini_JailBreakV_28K datasets will comprise the following columns:

  • id: Unique identifier for all samples.
  • jailbreak_query: Jailbreak_query obtained by different jailbreak attacks.
  • redteam_query: Harmful query from RedTeam_2K.
  • format: Jailbreak attack method including template, persuade, logic, figstep, query-relevant.
  • policy: The safety policy that redteam_query against.
  • image_path: The file path of the image.
  • from: The source of data.
  • selected_mini: "True" if the data in mini_JailBreakV_28K dataset, otherwise "False".
  • transfer_from_llm: "True" if the jailbreak_query is transferred from LLM jailbreak attacks, otherwise "False".

RedTeam_2K will comprise the following columns:

  • id: Unique identifier for all samples.
  • question: Harmful query.
  • policy: the safety policy that redteam_query against.
  • from: The source of data.

πŸš€ Data Composition

RedTeam-2K: RedTeam-2K dataset, a meticulously curated collection of 2, 000 harmful queries aimed at identifying alignment vulnerabilities within LLMs and MLLMs. This dataset spans across 16 safety policies and incorporates queries from 8 distinct sources. JailBreakV-28K: JailBreakV-28K contains 28, 000 jailbreak text-image pairs, which include 20, 000 text-based LLM transfer jailbreak attacks and 8, 000 image-based MLLM jailbreak attacks. This dataset covers 16 safety policies and 5 diverse jailbreak methods.

πŸ› οΈ Dataset Overview

The RedTeam-2K dataset, is a meticulously curated collection of 2, 000 harmful queries aimed at identifying alignment vulnerabilities within LLMs and MLLMs. This dataset spans 16 safety policies and incorporates queries from 8 distinct sources, including GPT Rewrite, Handcraft, GPT Generate, LLM Jailbreak Study, AdvBench, BeaverTails, Question Set, and hh-rlhf of Anthropic. Building upon the harmful query dataset provided by RedTeam-2K, JailBreakV-28K is designed as a comprehensive and diversified benchmark for evaluating the transferability of jailbreak attacks from LLMs to MLLMs, as well as assessing the alignment robustness of MLLMs against such attacks. Specifically, JailBreakV-28K contains 28, 000 jailbreak text-image pairs, which include 20, 000 text-based LLM transfer jailbreak attacks and 8, 000 image-based MLLM jailbreak attacks. This dataset covers 16 safety policies and 5 diverse jailbreak methods. The jailbreak methods are formed by 3 types of LLM transfer attacks that include Logic (Cognitive Overload), Persuade (Persuasive Adversarial Prompts), and Template (including both of Greedy Coordinate Gradient and handcrafted strategies), and 2 types of MLLM attacks including FigStep and Query-relevant attack. The JailBreakV-28K offers a broad spectrum of attack methodologies and integrates various image types like Nature, Random Noise, Typography, Stable Diffusion (SD), Blank, and SD+Typography Images. We believe JailBreakV-28K can serve as a comprehensive jailbreak benchmark for MLLMs.

πŸ† Mini-Leaderboard

ModelTotal ASRTransfer Attack ASR
OmniLMM-12B58.170.2
InfiMM-Zephyr-7B52.973.0
LLaMA-Adapter-v251.268.1
LLaVA-1.5-13B51.065.5
LLaVA-1.5-7B46.861.4
InstructBLIP-13B45.255.5
InternLM-XComposer2-VL-7B39.129.3
Bunny-v138.049.5
Qwen-VL-Chat33.741.2
InstructBLIP-7B26.046.8

❌ Disclaimers

This dataset contains offensive content that may be disturbing, This benchmark is provided for educational and research purposes only.

πŸ“² Contact

πŸ“– BibTeX:

@misc{luo2024jailbreakv28k,
      title={JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks}, 
      author={Weidi Luo and Siyuan Ma and Xiaogeng Liu and Xiaoyu Guo and Chaowei Xiao},
      year={2024},
      eprint={2404.03027},
      archivePrefix={arXiv},
      primaryClass={cs.CR}
}

[More Information Needed]

Contributors

EddyLuo

114 commits

EddyLuo/JailBreakV_28K

Dataset

69

stars

114

commits

2

linked in READMEs

Jul 10, 2024

updated

README

⛓‍πŸ’₯ JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

🌐 GitHub | πŸ›Ž Project Page | πŸ‘‰ Download full datasets

If you like our project, please give us a star ⭐ on Hugging Face for the latest update.

πŸ“° News

DateEvent
2024/07/09πŸŽ‰ Our paper is accepted by COLM 2024.
2024/06/22πŸ› οΈ We have updated our version to V0.2, which supports users to customize their attack models and evaluate models.
2024/04/04🎁 We have posted our paper on Arxiv.
2024/04/03πŸŽ‰ We have released our evaluation and inference samples.
2024/03/30πŸ”₯ We have released our dataset.

πŸ“₯ Using our dataset via huggingface Dataset

from datasets import load_dataset


mini_JailBreakV_28K = load_dataset("JailbreakV-28K/JailBreakV-28k", 'JailBreakV_28K')["mini_JailBreakV_28K"]
JailBreakV_28K = load_dataset("JailbreakV-28K/JailBreakV-28k", 'JailBreakV_28K')["JailBreakV_28K"]
RedTeam_2K = load_dataset("JailbreakV-28K/JailBreakV-28k", 'RedTeam_2K')["RedTeam_2K"]

πŸ‘» Inference and Evaluation

Create environment

conda create -n jbv python=3.9
conda activate jbv
pip install -r requirements.txt

Conduct jailbreak attack on MLLMs

# we default use Bunny-v1_0, you can change the default attack model to your customized attack models by editing the annotated codes.
# You can follow the Bunny script in <attack_models> to add other attack models.
python attack.py --root JailBreakV_28K 

Conduct evaluation

# we default use LlamaGuard, you can change the default evaluate model to your customized evaluate models by editing the annotated codes.
# You can follow the LlamaGuard script in <evaluate_models> to add other evaluate models.
python eval.py --data_path ./results/JailBreakV_28k/<your customized attack model>/JailBreakV_28K.csv

πŸ˜ƒ Dataset Details

JailBreakV_28K and mini_JailBreakV_28K datasets will comprise the following columns:

  • id: Unique identifier for all samples.
  • jailbreak_query: Jailbreak_query obtained by different jailbreak attacks.
  • redteam_query: Harmful query from RedTeam_2K.
  • format: Jailbreak attack method including template, persuade, logic, figstep, query-relevant.
  • policy: The safety policy that redteam_query against.
  • image_path: The file path of the image.
  • from: The source of data.
  • selected_mini: "True" if the data in mini_JailBreakV_28K dataset, otherwise "False".
  • transfer_from_llm: "True" if the jailbreak_query is transferred from LLM jailbreak attacks, otherwise "False".

RedTeam_2K will comprise the following columns:

  • id: Unique identifier for all samples.
  • question: Harmful query.
  • policy: the safety policy that redteam_query against.
  • from: The source of data.

πŸš€ Data Composition

RedTeam-2K: RedTeam-2K dataset, a meticulously curated collection of 2, 000 harmful queries aimed at identifying alignment vulnerabilities within LLMs and MLLMs. This dataset spans across 16 safety policies and incorporates queries from 8 distinct sources. JailBreakV-28K: JailBreakV-28K contains 28, 000 jailbreak text-image pairs, which include 20, 000 text-based LLM transfer jailbreak attacks and 8, 000 image-based MLLM jailbreak attacks. This dataset covers 16 safety policies and 5 diverse jailbreak methods.

πŸ› οΈ Dataset Overview

The RedTeam-2K dataset, is a meticulously curated collection of 2, 000 harmful queries aimed at identifying alignment vulnerabilities within LLMs and MLLMs. This dataset spans 16 safety policies and incorporates queries from 8 distinct sources, including GPT Rewrite, Handcraft, GPT Generate, LLM Jailbreak Study, AdvBench, BeaverTails, Question Set, and hh-rlhf of Anthropic. Building upon the harmful query dataset provided by RedTeam-2K, JailBreakV-28K is designed as a comprehensive and diversified benchmark for evaluating the transferability of jailbreak attacks from LLMs to MLLMs, as well as assessing the alignment robustness of MLLMs against such attacks. Specifically, JailBreakV-28K contains 28, 000 jailbreak text-image pairs, which include 20, 000 text-based LLM transfer jailbreak attacks and 8, 000 image-based MLLM jailbreak attacks. This dataset covers 16 safety policies and 5 diverse jailbreak methods. The jailbreak methods are formed by 3 types of LLM transfer attacks that include Logic (Cognitive Overload), Persuade (Persuasive Adversarial Prompts), and Template (including both of Greedy Coordinate Gradient and handcrafted strategies), and 2 types of MLLM attacks including FigStep and Query-relevant attack. The JailBreakV-28K offers a broad spectrum of attack methodologies and integrates various image types like Nature, Random Noise, Typography, Stable Diffusion (SD), Blank, and SD+Typography Images. We believe JailBreakV-28K can serve as a comprehensive jailbreak benchmark for MLLMs.

πŸ† Mini-Leaderboard

ModelTotal ASRTransfer Attack ASR
OmniLMM-12B58.170.2
InfiMM-Zephyr-7B52.973.0
LLaMA-Adapter-v251.268.1
LLaVA-1.5-13B51.065.5
LLaVA-1.5-7B46.861.4
InstructBLIP-13B45.255.5
InternLM-XComposer2-VL-7B39.129.3
Bunny-v138.049.5
Qwen-VL-Chat33.741.2
InstructBLIP-7B26.046.8

❌ Disclaimers

This dataset contains offensive content that may be disturbing, This benchmark is provided for educational and research purposes only.

πŸ“² Contact

πŸ“– BibTeX:

@misc{luo2024jailbreakv28k,
      title={JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks}, 
      author={Weidi Luo and Siyuan Ma and Xiaogeng Liu and Xiaoyu Guo and Chaowei Xiao},
      year={2024},
      eprint={2404.03027},
      archivePrefix={arXiv},
      primaryClass={cs.CR}
}

[More Information Needed]

Contributors

EddyLuo

114 commits