A mechanism-first classification of single-turn, inference-time prompt attacks on large language models, developed by the MLCommons AI Risk and Reliability (AIRR) working group.
MLCommons does not create or discover novel attacks. This taxonomy organizes and classifies jailbreak techniques that are already published in the academic literature, documented in practitioner reports, or submitted by external contributors through the open contribution process. Its purpose is to help defenders, researchers, and standards bodies understand the landscape of known attack mechanisms.
The taxonomy organizes known jailbreak mechanisms into four families based on the dominant prompt-manipulation strategy:
| Family | Prevalence | Categories | Leaves | Description |
|---|---|---|---|---|
| Perturbation | 36.28% | 2 | 4 | Modifies surface form while preserving semantic intent |
| Encoding Abuse | 20.35% | 2 | 5 | Manipulates representation or structural formatting |
| Overt Carriers | 19.47% | 2 | 4 | Applies explicit rhetorical pressure or narrative framing |
| Composition & Ordering | 23.89% | 2 | 5 | Arranges prompt structure to embed harmful objectives |
Total: 4 families, 8 categories, 18 leaf-level mechanisms.
See all 113 attacks in one place → Attacks Overview — a collapsible catalog organized by family, covering every attack documented in this taxonomy.
jailbreak-taxonomy/
├── README.md # This file
├── CHANGELOG.md # Version history
├── CITATION.cff # Machine-readable citation metadata
├── CONTRIBUTING.md # How to contribute
├── DATASHEET.md # Dataset documentation (Gebru et al. framework)
├── LICENSE # CC-BY-4.0
├── taxonomy/
│ ├── taxonomy.yaml # Machine-readable taxonomy structure
│ ├── attacks.yaml # Machine-readable attack catalog (113 attacks)
│ ├── taxonomy-overview.md # Human-readable summary with tree diagram
│ ├── attacks-overview.md # Auto-generated cross-family catalog (collapsible by family)
│ └── families/
│ ├── perturbation/
│ │ ├── README.md
│ │ ├── plain-perturbations.md
│ │ └── transfer-text-perturbations.md
│ ├── encoding-abuse/
│ │ ├── README.md
│ │ ├── encoding-unicode-tricks.md
│ │ └── wrappers-and-schemas.md
│ ├── overt-carriers/
│ │ ├── README.md
│ │ ├── direct-override-patterns.md
│ │ └── roleplay-and-template-jailbreaks.md
│ └── composition-ordering/
│ ├── README.md
│ ├── context-framing-and-deception.md
│ └── fragment-assembly.md
├── .github/
│ ├── ISSUE_TEMPLATE/ # GitHub issue templates (attack submission, taxonomy contribution, bug report)
│ ├── PULL_REQUEST_TEMPLATE.md # Default PR template
│ └── TAXONOMY_STRUCTURE_TEMPLATE.md # Detailed form for taxonomy structure proposals
├── governance/
│ ├── maintainers.md # Working group ownership, conflicts of interest, code of conduct, contact
│ └── review-process.md # Review pipeline and technical review criteria
└── releases/
└── v0.7.0/
└── taxonomy-snapshot.yaml
Attacks are grouped by how the prompt manipulates the model at inference time, not by the hazards they target or the outcomes they produce. A role-play attack and a Base64-encoded attack may both aim to extract the same harmful content, but they exploit fundamentally different mechanisms and require different defenses. The taxonomy classifies them separately for this reason.
The taxonomy uses a Family > Category > Leaf structure. Families capture the broadest mechanism distinction (surface perturbation vs. encoding manipulation vs. rhetorical override vs. compositional arrangement). Categories subdivide by a single explicit splitting rule. Leaves are the atomic units: each corresponds to a specific, testable manipulation strategy with documented instances in the literature.
This is a critical distinction. The taxonomy is the public universe of known jailbreak mechanisms, organized and documented for community use. The MLCommons Jailbreak Benchmark is a private selection from the taxonomy, instantiated with specific prompts and evaluated against specific systems under test.
All techniques in this taxonomy originate from published academic research, practitioner reports, or community submissions — MLCommons classifies and organizes these techniques but does not develop or discover them. Publishing the taxonomy does not reveal which techniques are tested in the benchmark. A model provider that defends against the entire taxonomy improves genuinely; one that tries to guess the benchmark's subset gains no reliable advantage. The taxonomy benefits defenders, researchers, and standards bodies. The benchmark's integrity is preserved by the separation.
The taxonomy satisfies six requirements from the v0.7 methodology (Section 5.2):
Researchers: Start with the taxonomy overview for the full tree, then read family and leaf documentation for mechanism details, inclusion/exclusion cues, canonical examples, and the attack catalog for each leaf. The attacks overview provides a single collapsible index of all 113 attacks grouped by family, useful for scanning the complete catalog in one place. The taxonomy.yaml provides the machine-readable taxonomy structure; attacks.yaml provides the machine-readable attack catalog with paper links, repositories, and model coverage for all 113 documented attacks.
Security teams and red-team practitioners: Use the leaf-level documentation to structure red-teaming exercises by mechanism family for coverage across structurally distinct attack strategies. The prevalence data helps prioritize testing.
Standards bodies and policymakers: Reference the taxonomy when specifying evaluation requirements or assessing benchmark coverage claims. The governance framework and versioning ensure longitudinal stability.
Contributors: See CONTRIBUTING.md for how to propose new mechanisms, place new attacks, or improve documentation.
The taxonomy uses semantic versioning:
Monthly stability releases produce a tagged snapshot with a stopping-rule check, prevalence update, and changelog entry. See CHANGELOG.md for version history.
The taxonomy is maintained by the Security Workstream of the MLCommons AI Risk and Reliability working group. All submissions are reviewed against the criteria documented in governance/review-process.md. The methodological basis is described in the v0.7 methodology paper.
See also:
This work is licensed under the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
If you use or reference this taxonomy, please cite the relevant papers:
@article{maple2026robust,
title={A Robust, Defensible, and Reproducible Methodology for Benchmarking
Single-Turn Jailbreak Attacks on Large Language Models},
author={Maple, Carsten and Tapwal, Riya and Knotz, Chris and Ezick, James
and Goel, James and Parrish, Alicia and Ricciuti, Federico and
Petit, Jonathan and Hui, Ong Chen and Lim, Seok Min and others},
year={2026},
month={February},
url={https://mlcommons.org/wp-content/uploads/2026/02/MLCommons___Security___Jailbreak_0_7_Paper_Collaborative.pdf}
}
@article{goel2025ailuminate,
title={AILuminate Security: Introducing v0.5 of the Jailbreak Benchmark
from MLCommons},
author={Goel, James and Parrish, Alicia and Knotz, Chris and Ezick, James
and Petit, Jonathan and Aroyo, Lora and Gruen, Andrew and
Bollacker, Kurt and Pietri, William and Hillenbrand, Bennett
and others},
year={2025},
month={October},
url={https://mlcommons.org/wp-content/uploads/2025/10/MLCommons___Security___Jailbreak_0_5_Paper-5.pdf}
}
@article{ghosh2025ailuminate,
title={AILuminate: Introducing v1.0 of the AI Safety Benchmark from
MLCommons},
author={Ghosh, Shaona and Parrish, Alicia and Bollacker, Kurt and
Goel, James and Pietri, William and Gruen, Andrew and
Aroyo, Lora and Knotz, Chris and Ezick, James and others},
year={2025},
eprint={2503.05731},
archivePrefix={arXiv},
url={https://arxiv.org/abs/2503.05731}
}
Contributions welcome. See CONTRIBUTING.md to get started.
Python
100.0%
A mechanism-first classification of single-turn, inference-time prompt attacks on large language models, developed by the MLCommons AI Risk and Reliability (AIRR) working group.
MLCommons does not create or discover novel attacks. This taxonomy organizes and classifies jailbreak techniques that are already published in the academic literature, documented in practitioner reports, or submitted by external contributors through the open contribution process. Its purpose is to help defenders, researchers, and standards bodies understand the landscape of known attack mechanisms.
The taxonomy organizes known jailbreak mechanisms into four families based on the dominant prompt-manipulation strategy:
| Family | Prevalence | Categories | Leaves | Description |
|---|---|---|---|---|
| Perturbation | 36.28% | 2 | 4 | Modifies surface form while preserving semantic intent |
| Encoding Abuse | 20.35% | 2 | 5 | Manipulates representation or structural formatting |
| Overt Carriers | 19.47% | 2 | 4 | Applies explicit rhetorical pressure or narrative framing |
| Composition & Ordering | 23.89% | 2 | 5 | Arranges prompt structure to embed harmful objectives |
Total: 4 families, 8 categories, 18 leaf-level mechanisms.
See all 113 attacks in one place → Attacks Overview — a collapsible catalog organized by family, covering every attack documented in this taxonomy.
jailbreak-taxonomy/
├── README.md # This file
├── CHANGELOG.md # Version history
├── CITATION.cff # Machine-readable citation metadata
├── CONTRIBUTING.md # How to contribute
├── DATASHEET.md # Dataset documentation (Gebru et al. framework)
├── LICENSE # CC-BY-4.0
├── taxonomy/
│ ├── taxonomy.yaml # Machine-readable taxonomy structure
│ ├── attacks.yaml # Machine-readable attack catalog (113 attacks)
│ ├── taxonomy-overview.md # Human-readable summary with tree diagram
│ ├── attacks-overview.md # Auto-generated cross-family catalog (collapsible by family)
│ └── families/
│ ├── perturbation/
│ │ ├── README.md
│ │ ├── plain-perturbations.md
│ │ └── transfer-text-perturbations.md
│ ├── encoding-abuse/
│ │ ├── README.md
│ │ ├── encoding-unicode-tricks.md
│ │ └── wrappers-and-schemas.md
│ ├── overt-carriers/
│ │ ├── README.md
│ │ ├── direct-override-patterns.md
│ │ └── roleplay-and-template-jailbreaks.md
│ └── composition-ordering/
│ ├── README.md
│ ├── context-framing-and-deception.md
│ └── fragment-assembly.md
├── .github/
│ ├── ISSUE_TEMPLATE/ # GitHub issue templates (attack submission, taxonomy contribution, bug report)
│ ├── PULL_REQUEST_TEMPLATE.md # Default PR template
│ └── TAXONOMY_STRUCTURE_TEMPLATE.md # Detailed form for taxonomy structure proposals
├── governance/
│ ├── maintainers.md # Working group ownership, conflicts of interest, code of conduct, contact
│ └── review-process.md # Review pipeline and technical review criteria
└── releases/
└── v0.7.0/
└── taxonomy-snapshot.yaml
Attacks are grouped by how the prompt manipulates the model at inference time, not by the hazards they target or the outcomes they produce. A role-play attack and a Base64-encoded attack may both aim to extract the same harmful content, but they exploit fundamentally different mechanisms and require different defenses. The taxonomy classifies them separately for this reason.
The taxonomy uses a Family > Category > Leaf structure. Families capture the broadest mechanism distinction (surface perturbation vs. encoding manipulation vs. rhetorical override vs. compositional arrangement). Categories subdivide by a single explicit splitting rule. Leaves are the atomic units: each corresponds to a specific, testable manipulation strategy with documented instances in the literature.
This is a critical distinction. The taxonomy is the public universe of known jailbreak mechanisms, organized and documented for community use. The MLCommons Jailbreak Benchmark is a private selection from the taxonomy, instantiated with specific prompts and evaluated against specific systems under test.
All techniques in this taxonomy originate from published academic research, practitioner reports, or community submissions — MLCommons classifies and organizes these techniques but does not develop or discover them. Publishing the taxonomy does not reveal which techniques are tested in the benchmark. A model provider that defends against the entire taxonomy improves genuinely; one that tries to guess the benchmark's subset gains no reliable advantage. The taxonomy benefits defenders, researchers, and standards bodies. The benchmark's integrity is preserved by the separation.
The taxonomy satisfies six requirements from the v0.7 methodology (Section 5.2):
Researchers: Start with the taxonomy overview for the full tree, then read family and leaf documentation for mechanism details, inclusion/exclusion cues, canonical examples, and the attack catalog for each leaf. The attacks overview provides a single collapsible index of all 113 attacks grouped by family, useful for scanning the complete catalog in one place. The taxonomy.yaml provides the machine-readable taxonomy structure; attacks.yaml provides the machine-readable attack catalog with paper links, repositories, and model coverage for all 113 documented attacks.
Security teams and red-team practitioners: Use the leaf-level documentation to structure red-teaming exercises by mechanism family for coverage across structurally distinct attack strategies. The prevalence data helps prioritize testing.
Standards bodies and policymakers: Reference the taxonomy when specifying evaluation requirements or assessing benchmark coverage claims. The governance framework and versioning ensure longitudinal stability.
Contributors: See CONTRIBUTING.md for how to propose new mechanisms, place new attacks, or improve documentation.
The taxonomy uses semantic versioning:
Monthly stability releases produce a tagged snapshot with a stopping-rule check, prevalence update, and changelog entry. See CHANGELOG.md for version history.
The taxonomy is maintained by the Security Workstream of the MLCommons AI Risk and Reliability working group. All submissions are reviewed against the criteria documented in governance/review-process.md. The methodological basis is described in the v0.7 methodology paper.
See also:
This work is licensed under the Creative Commons Attribution 4.0 International License (CC-BY-4.0).
If you use or reference this taxonomy, please cite the relevant papers:
@article{maple2026robust,
title={A Robust, Defensible, and Reproducible Methodology for Benchmarking
Single-Turn Jailbreak Attacks on Large Language Models},
author={Maple, Carsten and Tapwal, Riya and Knotz, Chris and Ezick, James
and Goel, James and Parrish, Alicia and Ricciuti, Federico and
Petit, Jonathan and Hui, Ong Chen and Lim, Seok Min and others},
year={2026},
month={February},
url={https://mlcommons.org/wp-content/uploads/2026/02/MLCommons___Security___Jailbreak_0_7_Paper_Collaborative.pdf}
}
@article{goel2025ailuminate,
title={AILuminate Security: Introducing v0.5 of the Jailbreak Benchmark
from MLCommons},
author={Goel, James and Parrish, Alicia and Knotz, Chris and Ezick, James
and Petit, Jonathan and Aroyo, Lora and Gruen, Andrew and
Bollacker, Kurt and Pietri, William and Hillenbrand, Bennett
and others},
year={2025},
month={October},
url={https://mlcommons.org/wp-content/uploads/2025/10/MLCommons___Security___Jailbreak_0_5_Paper-5.pdf}
}
@article{ghosh2025ailuminate,
title={AILuminate: Introducing v1.0 of the AI Safety Benchmark from
MLCommons},
author={Ghosh, Shaona and Parrish, Alicia and Bollacker, Kurt and
Goel, James and Pietri, William and Gruen, Andrew and
Aroyo, Lora and Knotz, Chris and Ezick, James and others},
year={2025},
eprint={2503.05731},
archivePrefix={arXiv},
url={https://arxiv.org/abs/2503.05731}
}
Contributions welcome. See CONTRIBUTING.md to get started.
Python
100.0%