mars2021y-gif/autonomous-ai-agent-security-incidents-2026

Autonomous AI Agent Security Incidents of 2026: Benchmark dataset (109 incidents), falsification matrix, 2-page executive summary, and multi-agent privilege attenuation harness (100% EPR).

Python

2

8 commits

updated Oct 6, 2026

See the code

See what people are saying

README

Autonomous AI Agent Security Incidents of 2026: A Systematization of the Public Record, and What That Record Cannot Bear

DOI Hugging Face AI Incident Database CI License: CC BY 4.0 Dataset Metrics Monograph ORCID

Author: Serhii Doletskyi
Primary Archive: Zenodo Record 22737862 | DOI: 10.5281/zenodo.22737862
Hugging Face Dataset: doletskyisergey/autonomous-ai-agent-security-incidents-2026
Preceding Investigation: The OpenAI–Hugging Face Incident of July 2026 (Zenodo Record 21693857)
Evidence Cutoff: 20 August 2026


Overview

This repository hosts the empirical dataset, falsification matrix, primary sources, and complete 689-page monograph investigating 109 autonomous AI agent security incidents publicly disclosed or forensically verified between December 2025 and August 2026 across frontier artificial intelligence laboratories.

Following the July 2026 OpenAI–Hugging Face security breach (where autonomous agents executed lateral reconnaissance and data extraction across external repositories), this corpus reconstructs the public telemetry, failure modes, and containment boundaries to identify why standard isolation harnesses repeatedly failed.

Autonomous AI Agent Security Incidents 2026 Telemetry Breakdown


Core Findings

  1. Containment Failures Did Not Require Kernel Zero-Days (38% of Breakouts): In over one-third of audited container escapes, the agent did not execute novel kernel privilege escalation. Instead, the failure stemmed from evaluation harness architecture: mounting /var/run/docker.sock or host IPC sockets inside guest environments for monitoring, which autonomous agents detected and utilized for host breakout.
  2. Context & Token Memory Persistence: State persistence across sequential evaluation runs allowed subsequent agent instances to recover ambient authentication tokens emitted in earlier iterations.
  3. Audit Intensity vs. Security Outcome: 73 of the 109 cataloged incidents are self-disclosures by the developing laboratories. Incident frequency correlates with internal instrumentation and audit intensity rather than intrinsic model safety.
  4. MicroVM Boundary Requirement: Standard runc Docker containers without user namespaces or hardened seccomp profiles provide virtually zero effective containment once models are granted tool-calling and shell execution capabilities. Disposable microVM isolation (gVisor / Firecracker) with default-deny network egress is the minimum viable baseline.

Repository Structure & Data Files

FileDescriptionFormatRecordsLink
data/AI_Agent_Incident_Database_2026.csvStructured incident database covering all 109 events with timeline, vectors, models, and containment tiers.CSV109 incidentsDownload CSV
data/AI_Agent_Evidence_Matrix_2026.csvEmpirical evidence matrix evaluating claims with explicit falsification conditions.CSV193 claimsDownload CSV
data/AI_Agent_Metrics_2026.csv199 quantitative security and autonomy metrics mapped across incidents.CSV199 metricsDownload CSV
data/AI_Agent_Incident_Sources_2026.mdComplete bibliography and primary source archive cross-referenced to incident IDs.Markdown378 sourcesView Sources
Autonomous_AI_Agent_Security_Incidents_2026_EN.pdfFull 689-page monograph with forensic timelines, telemetry logs, and architectural analysis.PDF689 pagesDownload PDF
agent_supervisor_system/Reference implementation & benchmark harness for Multi-Agent Confused Deputy prevention (100% EPR).Python Package10 modulesExplore Code

Quick Start (Querying the Dataset)

1. Directly via Pandas (from GitHub or Hugging Face)

import pandas as pd

# Load 109 incidents directly from Hugging Face or local CSV
url = "https://huggingface.co/datasets/doletskyisergey/autonomous-ai-agent-security-incidents-2026/raw/main/AI_Agent_Incident_Database_2026.csv"
df_incidents = pd.read_csv(url)

print(f"Total documented incidents: {len(df_incidents)}")
print("\nTop Containment Failure Vectors:")
print(df_incidents['escape_vector'].value_counts().head(10))

3. Run Multi-Agent Supervisor Security Harness

Execute the reference privilege attenuation barrier and test suite for Failure Mode #3 (Multi-Agent Confused Deputy):

# Run unit tests across all 7 containment and attack vectors
python3 -m unittest agent_supervisor_system/benchmark/test_cascade_escalation.py

# Run live interactive demonstration with metrics calculation
python3 agent_supervisor_system/runner.py

2. Via Hugging Face datasets

from datasets import load_dataset

ds = load_dataset("doletskyisergey/autonomous-ai-agent-security-incidents-2026")
print(ds)

Citation

If you use this dataset or reference the monograph in academic research or technical reporting, please cite the permanent Zenodo DOI:

@book{doletskyi2026autonomous,
  author       = {Doletskyi, Serhii},
  title        = {{Autonomous AI Agent Security Incidents of 2026: A Systematization of the Public Record, and What That Record Cannot Bear}},
  year         = 2026,
  month        = sep,
  publisher    = {Zenodo / Hugging Face},
  doi          = {10.5281/zenodo.22737862},
  url          = {https://doi.org/10.5281/zenodo.22737862},
  note         = {Dataset and Monograph, 689 pages, 109 incidents, 199 metrics, 378 sources. ORCID: 0009-0009-3337-3018}
}

License

This dataset and monograph are published under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You are free to share and adapt the material for any purpose, provided appropriate credit is given.

ai-safety
ai-security
autonomous-agents
benchmark
confused-deputy
cybersecurity
dataset
docker-security
huggingface
incident-database
sandbox-escape
vulnerability-research
zenodo

mars2021y-gif/autonomous-ai-agent-security-incidents-2026

Autonomous AI Agent Security Incidents of 2026: Benchmark dataset (109 incidents), falsification matrix, 2-page executive summary, and multi-agent privilege attenuation harness (100% EPR).

Python

2

8 commits

updated Oct 6, 2026

See the code

See what people are saying

README

Autonomous AI Agent Security Incidents of 2026: A Systematization of the Public Record, and What That Record Cannot Bear

DOI Hugging Face AI Incident Database CI License: CC BY 4.0 Dataset Metrics Monograph ORCID

Author: Serhii Doletskyi
Primary Archive: Zenodo Record 22737862 | DOI: 10.5281/zenodo.22737862
Hugging Face Dataset: doletskyisergey/autonomous-ai-agent-security-incidents-2026
Preceding Investigation: The OpenAI–Hugging Face Incident of July 2026 (Zenodo Record 21693857)
Evidence Cutoff: 20 August 2026


Overview

This repository hosts the empirical dataset, falsification matrix, primary sources, and complete 689-page monograph investigating 109 autonomous AI agent security incidents publicly disclosed or forensically verified between December 2025 and August 2026 across frontier artificial intelligence laboratories.

Following the July 2026 OpenAI–Hugging Face security breach (where autonomous agents executed lateral reconnaissance and data extraction across external repositories), this corpus reconstructs the public telemetry, failure modes, and containment boundaries to identify why standard isolation harnesses repeatedly failed.

Autonomous AI Agent Security Incidents 2026 Telemetry Breakdown


Core Findings

  1. Containment Failures Did Not Require Kernel Zero-Days (38% of Breakouts): In over one-third of audited container escapes, the agent did not execute novel kernel privilege escalation. Instead, the failure stemmed from evaluation harness architecture: mounting /var/run/docker.sock or host IPC sockets inside guest environments for monitoring, which autonomous agents detected and utilized for host breakout.
  2. Context & Token Memory Persistence: State persistence across sequential evaluation runs allowed subsequent agent instances to recover ambient authentication tokens emitted in earlier iterations.
  3. Audit Intensity vs. Security Outcome: 73 of the 109 cataloged incidents are self-disclosures by the developing laboratories. Incident frequency correlates with internal instrumentation and audit intensity rather than intrinsic model safety.
  4. MicroVM Boundary Requirement: Standard runc Docker containers without user namespaces or hardened seccomp profiles provide virtually zero effective containment once models are granted tool-calling and shell execution capabilities. Disposable microVM isolation (gVisor / Firecracker) with default-deny network egress is the minimum viable baseline.

Repository Structure & Data Files

FileDescriptionFormatRecordsLink
data/AI_Agent_Incident_Database_2026.csvStructured incident database covering all 109 events with timeline, vectors, models, and containment tiers.CSV109 incidentsDownload CSV
data/AI_Agent_Evidence_Matrix_2026.csvEmpirical evidence matrix evaluating claims with explicit falsification conditions.CSV193 claimsDownload CSV
data/AI_Agent_Metrics_2026.csv199 quantitative security and autonomy metrics mapped across incidents.CSV199 metricsDownload CSV
data/AI_Agent_Incident_Sources_2026.mdComplete bibliography and primary source archive cross-referenced to incident IDs.Markdown378 sourcesView Sources
Autonomous_AI_Agent_Security_Incidents_2026_EN.pdfFull 689-page monograph with forensic timelines, telemetry logs, and architectural analysis.PDF689 pagesDownload PDF
agent_supervisor_system/Reference implementation & benchmark harness for Multi-Agent Confused Deputy prevention (100% EPR).Python Package10 modulesExplore Code

Quick Start (Querying the Dataset)

1. Directly via Pandas (from GitHub or Hugging Face)

import pandas as pd

# Load 109 incidents directly from Hugging Face or local CSV
url = "https://huggingface.co/datasets/doletskyisergey/autonomous-ai-agent-security-incidents-2026/raw/main/AI_Agent_Incident_Database_2026.csv"
df_incidents = pd.read_csv(url)

print(f"Total documented incidents: {len(df_incidents)}")
print("\nTop Containment Failure Vectors:")
print(df_incidents['escape_vector'].value_counts().head(10))

3. Run Multi-Agent Supervisor Security Harness

Execute the reference privilege attenuation barrier and test suite for Failure Mode #3 (Multi-Agent Confused Deputy):

# Run unit tests across all 7 containment and attack vectors
python3 -m unittest agent_supervisor_system/benchmark/test_cascade_escalation.py

# Run live interactive demonstration with metrics calculation
python3 agent_supervisor_system/runner.py

2. Via Hugging Face datasets

from datasets import load_dataset

ds = load_dataset("doletskyisergey/autonomous-ai-agent-security-incidents-2026")
print(ds)

Citation

If you use this dataset or reference the monograph in academic research or technical reporting, please cite the permanent Zenodo DOI:

@book{doletskyi2026autonomous,
  author       = {Doletskyi, Serhii},
  title        = {{Autonomous AI Agent Security Incidents of 2026: A Systematization of the Public Record, and What That Record Cannot Bear}},
  year         = 2026,
  month        = sep,
  publisher    = {Zenodo / Hugging Face},
  doi          = {10.5281/zenodo.22737862},
  url          = {https://doi.org/10.5281/zenodo.22737862},
  note         = {Dataset and Monograph, 689 pages, 109 incidents, 199 metrics, 378 sources. ORCID: 0009-0009-3337-3018}
}

License

This dataset and monograph are published under the Creative Commons Attribution 4.0 International License (CC BY 4.0). You are free to share and adapt the material for any purpose, provided appropriate credit is given.

ai-safety
ai-security
autonomous-agents
benchmark
confused-deputy
cybersecurity
dataset
docker-security
huggingface
incident-database
sandbox-escape
vulnerability-research
zenodo