Curated AI security and safety evaluation benchmarks well-regarded by Frontier AI labs
See the codeA curated, categorized list of AI security and safety evaluation benchmarks well-regarded by Frontier AI labs (Anthropic, OpenAI, Google DeepMind, Meta) and AI Safety Institutes (US AISI, UK AISI).
Maintained by Anshu Gupta, Founder & CISO, Fixin Security. Founder, Tejas Cyber Network
| Category | Count |
|---|---|
| Cyber Offense and CTF | 9 |
| Cyber Defense and Threat Intel | 8 |
| Software Security and Code | 2 |
| Agent Security and Prompt Injection | 8 |
| Jailbreak and Refusal | 6 |
| CBRN Knowledge and Bio Uplift | 9 |
| Alignment, Honesty, Scheming | 4 |
| Autonomy and AI R&D | 4 |
| Comprehensive Safety and Trust | 6 |
| Bias and Fairness | 1 |
| Tooling Frameworks | 4 |
Public cyber capabilities benchmark of 40 CTF challenges from four CTF competitions
Targeted vulnerability reproduction in real open-source projects from high-level descriptions
Identify and exploit vulnerabilities in free and open-source web applications
15 cyber offense challenges aligned to MITRE ATT&CK, with 80 elicitation configurations to find best-performing setup
Difficulty scoring system for vulnerability and exploit benchmarks
Scenario-based benchmarking for LLM cyber capabilities
200+ CTF challenges from NYU CSAW competitions; complements Cybench
Automated vulnerability detection on Nginx and DARPA AIxCC framework
Open-source applications frozen at vulnerable versions, measuring miss rate on known CVEs
End-to-end detection rule generation with AI agents
Evaluating LLM agents on cyber threat investigation
Malware analysis and threat intelligence reasoning; defensive capabilities benchmark
MCQA, RCM, VSP, ATE tasks for cyber threat intelligence (knowledge, attribution, severity)
RAG-based benchmark for cybersecurity knowledge (cryptography, reverse engineering, risk)
Multi-dimensional cybersecurity benchmark: 44,823 MCQs and 3,087 SAQs across sub-domains
MCQs across software, network, and web security topics
Foundational cybersecurity concept questions
Security-oriented software engineering benchmark
Umbrella suite: insecure coding (CWE), MITRE ATT&CK helpfulness, prompt injection (textual and visual), code interpreter abuse, and CyberSOCEval
Curated set of high-impact attacks from large-scale public competition
Evaluating sabotage and monitoring in LLM agents (29 complex environments)
Dynamic framework jointly evaluating utility and prompt injection resilience for tool-integrated agents
Indirect prompt injection: 1,054 test cases, 17 user tools, 62 attacker tools
Benchmark for measuring harmfulness of LLM agents when user is malicious (ICLR 2025)
Benchmark for Indirect Prompt Injection Attacks
Prompt extraction and hijacking benchmark grown from a public game
Browser agent red teaming benchmark
State-of-the-art LLM jailbreak evaluation benchmark with quality-aware scoring
Standardized red-teaming evaluation framework with classifier-based harm grading
Open robustness benchmark for jailbreaking LLMs (NeurIPS 2024)
Tests over-refusal: incorrectly refusing safe requests (counterweight to StrongREJECT)
Fine-grained refusal evaluation across 45 unsafe topic categories
Adversarial harmful behaviors dataset (Zou et al. GCG paper)
3,668 MCQs across biosecurity, cybersecurity, and chemical security; proxy for hazardous knowledge and unlearning benchmark
Biology research capability: LitQA2, ProtocolQA, SeqQA, FigQA, Cloning Scenarios
Multiple-response virology benchmark; top models now exceed expert virologists
Long-form biorisk question evaluation
Bio tacit knowledge and troubleshooting questions
Creative biology task evaluations
Short-horizon computational biology tasks
WMD proliferation risk benchmark with safety-usefulness tradeoff
Real-world risk metric layered on top of LAB-Bench, BioLP-bench, and WMDP
Disentangles honesty from accuracy; large-scale lying-under-pressure evaluation
Six agentic evaluations where models are placed in environments that incentivize scheming
11 evaluations supporting a scheming-inability safety case
Tests model self-awareness as a propensity benchmark
AI R&D capabilities of language model agents vs human experts; multi-hour task time horizons
Task-length-AI-can-complete methodology; exponential trend tracking
Autonomous ML research task benchmark
Human-verified subset of real GitHub issues; used as autonomy signal in RSP and Preparedness contexts
12 hazard categories (violent crime, CSAM, weapons, suicide, privacy, defamation, hate, etc.); MLCommons industry standard
Comprehensive AI risk taxonomy benchmark spanning multiple safety dimensions
8 trustworthiness perspectives: toxicity, bias, robustness, privacy, ethics, fairness, OOD, adversarial
30+ datasets across 6 trust dimensions (truthfulness, safety, fairness, robustness, privacy, ethics)
11,000+ MCQs across 7 safety categories
Aggregator of 35+ safety benchmarks
Hand-built bias benchmark across nine demographic axes for QA
Evaluation harness used by US and UK AI Safety Institutes; AgentDojo and many others ship as Inspect tasks
Adversarial Threat Landscape for AI Systems; threat-model taxonomy (not a benchmark)
Python Risk Identification Toolkit; open-source red-teaming framework
Open-source LLM vulnerability scanner
If you want the tightest core list, these appear most consistently in 2025-2026 system cards from Anthropic, OpenAI, Google DeepMind, and Meta, plus AISI publications:
Pull requests welcome. Please include the paper URL, the publishing organization, and which frontier labs or AISIs have cited the benchmark.
This list is shared under CC BY 4.0. Linked papers and repositories retain their own licenses.
1 commits
Curated AI security and safety evaluation benchmarks well-regarded by Frontier AI labs
See the codeA curated, categorized list of AI security and safety evaluation benchmarks well-regarded by Frontier AI labs (Anthropic, OpenAI, Google DeepMind, Meta) and AI Safety Institutes (US AISI, UK AISI).
Maintained by Anshu Gupta, Founder & CISO, Fixin Security. Founder, Tejas Cyber Network
| Category | Count |
|---|---|
| Cyber Offense and CTF | 9 |
| Cyber Defense and Threat Intel | 8 |
| Software Security and Code | 2 |
| Agent Security and Prompt Injection | 8 |
| Jailbreak and Refusal | 6 |
| CBRN Knowledge and Bio Uplift | 9 |
| Alignment, Honesty, Scheming | 4 |
| Autonomy and AI R&D | 4 |
| Comprehensive Safety and Trust | 6 |
| Bias and Fairness | 1 |
| Tooling Frameworks | 4 |
Public cyber capabilities benchmark of 40 CTF challenges from four CTF competitions
Targeted vulnerability reproduction in real open-source projects from high-level descriptions
Identify and exploit vulnerabilities in free and open-source web applications
15 cyber offense challenges aligned to MITRE ATT&CK, with 80 elicitation configurations to find best-performing setup
Difficulty scoring system for vulnerability and exploit benchmarks
Scenario-based benchmarking for LLM cyber capabilities
200+ CTF challenges from NYU CSAW competitions; complements Cybench
Automated vulnerability detection on Nginx and DARPA AIxCC framework
Open-source applications frozen at vulnerable versions, measuring miss rate on known CVEs
End-to-end detection rule generation with AI agents
Evaluating LLM agents on cyber threat investigation
Malware analysis and threat intelligence reasoning; defensive capabilities benchmark
MCQA, RCM, VSP, ATE tasks for cyber threat intelligence (knowledge, attribution, severity)
RAG-based benchmark for cybersecurity knowledge (cryptography, reverse engineering, risk)
Multi-dimensional cybersecurity benchmark: 44,823 MCQs and 3,087 SAQs across sub-domains
MCQs across software, network, and web security topics
Foundational cybersecurity concept questions
Security-oriented software engineering benchmark
Umbrella suite: insecure coding (CWE), MITRE ATT&CK helpfulness, prompt injection (textual and visual), code interpreter abuse, and CyberSOCEval
Curated set of high-impact attacks from large-scale public competition
Evaluating sabotage and monitoring in LLM agents (29 complex environments)
Dynamic framework jointly evaluating utility and prompt injection resilience for tool-integrated agents
Indirect prompt injection: 1,054 test cases, 17 user tools, 62 attacker tools
Benchmark for measuring harmfulness of LLM agents when user is malicious (ICLR 2025)
Benchmark for Indirect Prompt Injection Attacks
Prompt extraction and hijacking benchmark grown from a public game
Browser agent red teaming benchmark
State-of-the-art LLM jailbreak evaluation benchmark with quality-aware scoring
Standardized red-teaming evaluation framework with classifier-based harm grading
Open robustness benchmark for jailbreaking LLMs (NeurIPS 2024)
Tests over-refusal: incorrectly refusing safe requests (counterweight to StrongREJECT)
Fine-grained refusal evaluation across 45 unsafe topic categories
Adversarial harmful behaviors dataset (Zou et al. GCG paper)
3,668 MCQs across biosecurity, cybersecurity, and chemical security; proxy for hazardous knowledge and unlearning benchmark
Biology research capability: LitQA2, ProtocolQA, SeqQA, FigQA, Cloning Scenarios
Multiple-response virology benchmark; top models now exceed expert virologists
Long-form biorisk question evaluation
Bio tacit knowledge and troubleshooting questions
Creative biology task evaluations
Short-horizon computational biology tasks
WMD proliferation risk benchmark with safety-usefulness tradeoff
Real-world risk metric layered on top of LAB-Bench, BioLP-bench, and WMDP
Disentangles honesty from accuracy; large-scale lying-under-pressure evaluation
Six agentic evaluations where models are placed in environments that incentivize scheming
11 evaluations supporting a scheming-inability safety case
Tests model self-awareness as a propensity benchmark
AI R&D capabilities of language model agents vs human experts; multi-hour task time horizons
Task-length-AI-can-complete methodology; exponential trend tracking
Autonomous ML research task benchmark
Human-verified subset of real GitHub issues; used as autonomy signal in RSP and Preparedness contexts
12 hazard categories (violent crime, CSAM, weapons, suicide, privacy, defamation, hate, etc.); MLCommons industry standard
Comprehensive AI risk taxonomy benchmark spanning multiple safety dimensions
8 trustworthiness perspectives: toxicity, bias, robustness, privacy, ethics, fairness, OOD, adversarial
30+ datasets across 6 trust dimensions (truthfulness, safety, fairness, robustness, privacy, ethics)
11,000+ MCQs across 7 safety categories
Aggregator of 35+ safety benchmarks
Hand-built bias benchmark across nine demographic axes for QA
Evaluation harness used by US and UK AI Safety Institutes; AgentDojo and many others ship as Inspect tasks
Adversarial Threat Landscape for AI Systems; threat-model taxonomy (not a benchmark)
Python Risk Identification Toolkit; open-source red-teaming framework
Open-source LLM vulnerability scanner
If you want the tightest core list, these appear most consistently in 2025-2026 system cards from Anthropic, OpenAI, Google DeepMind, and Meta, plus AISI publications:
Pull requests welcome. Please include the paper URL, the publishing organization, and which frontier labs or AISIs have cited the benchmark.
This list is shared under CC BY 4.0. Linked papers and repositories retain their own licenses.
1 commits