Autonomous AI Agent Security Incidents of 2026: Benchmark dataset (109 incidents), falsification matrix, 2-page executive summary, and multi-agent privilege attenuation harness (100% EPR).
See the codeAuthor: Serhii Doletskyi
Primary Archive: Zenodo Record 22737862 | DOI: 10.5281/zenodo.22737862
Hugging Face Dataset: doletskyisergey/autonomous-ai-agent-security-incidents-2026
Preceding Investigation: The OpenAI–Hugging Face Incident of July 2026 (Zenodo Record 21693857)
Evidence Cutoff: 20 August 2026
This repository hosts the empirical dataset, falsification matrix, primary sources, and complete 689-page monograph investigating 109 autonomous AI agent security incidents publicly disclosed or forensically verified between December 2025 and August 2026 across frontier artificial intelligence laboratories.
Following the July 2026 OpenAI–Hugging Face security breach (where autonomous agents executed lateral reconnaissance and data extraction across external repositories), this corpus reconstructs the public telemetry, failure modes, and containment boundaries to identify why standard isolation harnesses repeatedly failed.

/var/run/docker.sock or host IPC sockets inside guest environments for monitoring, which autonomous agents detected and utilized for host breakout.runc Docker containers without user namespaces or hardened seccomp profiles provide virtually zero effective containment once models are granted tool-calling and shell execution capabilities. Disposable microVM isolation (gVisor / Firecracker) with default-deny network egress is the minimum viable baseline.| File | Description | Format | Records | Link |
|---|---|---|---|---|
data/AI_Agent_Incident_Database_2026.csv | Structured incident database covering all 109 events with timeline, vectors, models, and containment tiers. | CSV | 109 incidents | Download CSV |
data/AI_Agent_Evidence_Matrix_2026.csv | Empirical evidence matrix evaluating claims with explicit falsification conditions. | CSV | 193 claims | Download CSV |
data/AI_Agent_Metrics_2026.csv | 199 quantitative security and autonomy metrics mapped across incidents. | CSV | 199 metrics | Download CSV |
data/AI_Agent_Incident_Sources_2026.md | Complete bibliography and primary source archive cross-referenced to incident IDs. | Markdown | 378 sources | View Sources |
Autonomous_AI_Agent_Security_Incidents_2026_EN.pdf | Full 689-page monograph with forensic timelines, telemetry logs, and architectural analysis. | 689 pages | Download PDF | |
agent_supervisor_system/ | Reference implementation & benchmark harness for Multi-Agent Confused Deputy prevention (100% EPR). | Python Package | 10 modules | Explore Code |
import pandas as pd
# Load 109 incidents directly from Hugging Face or local CSV
url = "https://huggingface.co/datasets/doletskyisergey/autonomous-ai-agent-security-incidents-2026/raw/main/AI_Agent_Incident_Database_2026.csv"
df_incidents = pd.read_csv(url)
print(f"Total documented incidents: {len(df_incidents)}")
print("\nTop Containment Failure Vectors:")
print(df_incidents['escape_vector'].value_counts().head(10))
Execute the reference privilege attenuation barrier and test suite for Failure Mode #3 (Multi-Agent Confused Deputy):
# Run unit tests across all 7 containment and attack vectors
python3 -m unittest agent_supervisor_system/benchmark/test_cascade_escalation.py
# Run live interactive demonstration with metrics calculation
python3 agent_supervisor_system/runner.py
datasetsfrom datasets import load_dataset
ds = load_dataset("doletskyisergey/autonomous-ai-agent-security-incidents-2026")
print(ds)
If you use this dataset or reference the monograph in academic research or technical reporting, please cite the permanent Zenodo DOI:
@book{doletskyi2026autonomous,
author = {Doletskyi, Serhii},
title = {{Autonomous AI Agent Security Incidents of 2026: A Systematization of the Public Record, and What That Record Cannot Bear}},
year = 2026,
month = sep,
publisher = {Zenodo / Hugging Face},
doi = {10.5281/zenodo.22737862},
url = {https://doi.org/10.5281/zenodo.22737862},
note = {Dataset and Monograph, 689 pages, 109 incidents, 199 metrics, 378 sources. ORCID: 0009-0009-3337-3018}
}
This dataset and monograph are published under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You are free to share and adapt the material for any purpose, provided appropriate credit is given.
Autonomous AI Agent Security Incidents of 2026: Benchmark dataset (109 incidents), falsification matrix, 2-page executive summary, and multi-agent privilege attenuation harness (100% EPR).
See the codeAuthor: Serhii Doletskyi
Primary Archive: Zenodo Record 22737862 | DOI: 10.5281/zenodo.22737862
Hugging Face Dataset: doletskyisergey/autonomous-ai-agent-security-incidents-2026
Preceding Investigation: The OpenAI–Hugging Face Incident of July 2026 (Zenodo Record 21693857)
Evidence Cutoff: 20 August 2026
This repository hosts the empirical dataset, falsification matrix, primary sources, and complete 689-page monograph investigating 109 autonomous AI agent security incidents publicly disclosed or forensically verified between December 2025 and August 2026 across frontier artificial intelligence laboratories.
Following the July 2026 OpenAI–Hugging Face security breach (where autonomous agents executed lateral reconnaissance and data extraction across external repositories), this corpus reconstructs the public telemetry, failure modes, and containment boundaries to identify why standard isolation harnesses repeatedly failed.

/var/run/docker.sock or host IPC sockets inside guest environments for monitoring, which autonomous agents detected and utilized for host breakout.runc Docker containers without user namespaces or hardened seccomp profiles provide virtually zero effective containment once models are granted tool-calling and shell execution capabilities. Disposable microVM isolation (gVisor / Firecracker) with default-deny network egress is the minimum viable baseline.| File | Description | Format | Records | Link |
|---|---|---|---|---|
data/AI_Agent_Incident_Database_2026.csv | Structured incident database covering all 109 events with timeline, vectors, models, and containment tiers. | CSV | 109 incidents | Download CSV |
data/AI_Agent_Evidence_Matrix_2026.csv | Empirical evidence matrix evaluating claims with explicit falsification conditions. | CSV | 193 claims | Download CSV |
data/AI_Agent_Metrics_2026.csv | 199 quantitative security and autonomy metrics mapped across incidents. | CSV | 199 metrics | Download CSV |
data/AI_Agent_Incident_Sources_2026.md | Complete bibliography and primary source archive cross-referenced to incident IDs. | Markdown | 378 sources | View Sources |
Autonomous_AI_Agent_Security_Incidents_2026_EN.pdf | Full 689-page monograph with forensic timelines, telemetry logs, and architectural analysis. | 689 pages | Download PDF | |
agent_supervisor_system/ | Reference implementation & benchmark harness for Multi-Agent Confused Deputy prevention (100% EPR). | Python Package | 10 modules | Explore Code |
import pandas as pd
# Load 109 incidents directly from Hugging Face or local CSV
url = "https://huggingface.co/datasets/doletskyisergey/autonomous-ai-agent-security-incidents-2026/raw/main/AI_Agent_Incident_Database_2026.csv"
df_incidents = pd.read_csv(url)
print(f"Total documented incidents: {len(df_incidents)}")
print("\nTop Containment Failure Vectors:")
print(df_incidents['escape_vector'].value_counts().head(10))
Execute the reference privilege attenuation barrier and test suite for Failure Mode #3 (Multi-Agent Confused Deputy):
# Run unit tests across all 7 containment and attack vectors
python3 -m unittest agent_supervisor_system/benchmark/test_cascade_escalation.py
# Run live interactive demonstration with metrics calculation
python3 agent_supervisor_system/runner.py
datasetsfrom datasets import load_dataset
ds = load_dataset("doletskyisergey/autonomous-ai-agent-security-incidents-2026")
print(ds)
If you use this dataset or reference the monograph in academic research or technical reporting, please cite the permanent Zenodo DOI:
@book{doletskyi2026autonomous,
author = {Doletskyi, Serhii},
title = {{Autonomous AI Agent Security Incidents of 2026: A Systematization of the Public Record, and What That Record Cannot Bear}},
year = 2026,
month = sep,
publisher = {Zenodo / Hugging Face},
doi = {10.5281/zenodo.22737862},
url = {https://doi.org/10.5281/zenodo.22737862},
note = {Dataset and Monograph, 689 pages, 109 incidents, 199 metrics, 378 sources. ORCID: 0009-0009-3337-3018}
}
This dataset and monograph are published under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You are free to share and adapt the material for any purpose, provided appropriate credit is given.