11 repos
xunguangwang/JailbreakGuardrailBenchmark
An Open Benchmark for Evaluating Jailbreak Guardrails in Large Language Models
5
5 commits
JailbreakBench/JBB-Behaviors
JailbreakBench
126
12 commits
EddyLuo/JailBreakV_28K
⛓💥 JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models…
70
114 commits
SaFo-Lab/JailBreakV_28K
[COLM 2024] JailBreakV-28K: A comprehensive benchmark designed to evaluate the transferability of…
98
41 commits
EddyLuo1232/JailBreakV_28K
SafeMTData/SafeMTData
💥Derail Yourself: Multi-turn LLM Jailbreak Attack through Self-discovered Clues
14
15 commits
allenai/wildteaming
No description
43
151 commits
JailbreakBench/jailbreakbench
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models [NeurIPS 2024…
673
269 commits
verazuo/jailbreak_llms
[CCS'24] A dataset consists of 15,140 ChatGPT prompts from Reddit, Discord, websites, and…
3,816
19 commits
Simsonsun/JailbreakPrompts
Independent Jailbreak Datasets for LLM Guardrail Evaluation
10
29 commits
xunguangwang/SoK4JailbreakGuardrails
[S&P 2026] SoK: Evaluating Jailbreak Guardrails for Large Language Models
48
13 commits