Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks

With the continuous evolution of Large Language Models (LLMs), LLM-based agents have advanced beyond passive chatbots to become autonomous cyber entities capable of performing complex tasks, including web browsing, malicious code and deceptive content generation, and decision-making. By significantly reducing the time, expertise, and resources, AI-assisted cyberattacks orchestrated by LLM-based agents have led to a phenomenon termed Cyber Threat Inflation, characterized by a significant reduction in attack costs and a tremendous increase in attack scale. To provide actionable defensive insights, in this survey, we focus on the potential cyber threats posed by LLM-based agents across diverse network systems. Firstly, we present the capabilities of LLM-based cyberattack agents, which include executing autonomous attack strategies, comprising scouting, memory, reasoning, and action, and facilitating collaborative operations with other agents or human operators. Building on these capabilities, we examine common cyberattacks initiated by LLM-based agents and compare their effectiveness across different types of networks, including static, mobile, and infrastructure-free paradigms. Moreover, we analyze threat bottlenecks of LLM-based agents across different network infrastructures and review their defense methods. Due to operational imbalances, existing defense methods are inadequate against autonomous cyberattacks. Finally, we outline future research directions and potential defensive strategies for legacy network systems.
Paper Link (arXiv)
Table of Contents
II. Large Language Model-based Agents for Autonomous Cyberattacks

Models
- Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks | Paper Link
- The Best Defense is a Good Offense: Countering LLM-Powered Cyberattacks | Paper Link
- BadAgent: Inserting and Activating Backdoor Attacks in LLM Agents | Paper Link
- LLMs Killed the Script Kiddie: How Agents Supported by Large Language Models Change the Landscape of Network Threat Testing | Paper Link
- Emerging Cyber Attack Risks of Medical AI Agents | Paper Link
- Hackphyr: A Local Fine-Tuned LLM Agent for Network Security Environments | Paper Link
- AttackLLM: LLM-based Attack Pattern Generation for an Industrial Control System | Paper Link
- Evaluating Frontier Models for Dangerous Capabilities | Paper Link
- Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMs | Paper Link
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents | Paper Link
- CS-Eval: A Comprehensive Large Language Model Benchmark for CyberSecurity | Paper Link
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal | Paper Link
- R-Judge: Benchmarking Safety Risk Awareness for LLM Agents | Paper Link
Perception
- When LLMs Meet Cybersecurity: A Systematic Literature Review | Paper Link
- Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness | Paper Link
- CyberPal.AI: Empowering LLMs with Expert-Driven Cybersecurity Instructions | Paper Link
A2. Memory
- A Survey on Large Language Model based Autonomous Agents | Paper Link
- Large Language Model-Based Agents for Software Engineering: A Survey | Paper Link
- The Rise and Potential of Large Language Model Based Agents: A Survey | Paper Link
- Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training | Paper Link
- AttackER: Towards Enhancing Cyber-Attack Attribution with a Named Entity Recognition Dataset | Paper Link
- SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security | Paper Link
- CmdCaliper: A Semantic-Aware Command-Line Embedding Model and Dataset for Security Research | Paper Link
- Retrieval-Augmented Generation for Large Language Models: A Survey | Paper Link
- Unifying Large Language Models and Knowledge Graphs: A Roadmap | Paper Link
- Exploring RAG-based Vulnerability Augmentation with LLMs | Paper Link
- AttacKG+: Boosting Attack Knowledge Graph Construction with Large Language Models | Paper Link
- CTIKG: LLM-Powered Knowledge Graph Construction from Cyber Threat Intelligence | Paper Link
- CTINexus: Automatic Cyber Threat Intelligence Knowledge Graph Construction Using Large Language Models | Paper Link
A3. Reasoning and Planning
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models | Paper Link
- Building Cyber Attack Trees with the Help of My LLM? A Mixed Method Study | Paper Link
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models | Paper Link
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models | Paper Link
- From Sands to Mansions: Enabling Automatic Full-Life-Cycle Cyberattack Construction with LLM | Paper Link
- ReAct: Synergizing Reasoning and Acting in Language Models | Paper Link
- LLM-Assisted Proactive Threat Intelligence for Automated Reasoning | Paper Link
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena | Paper Link
- Hackphyr: A Local Fine-Tuned LLM Agent for Network Security Environments | Paper Link
- Crimson: Empowering Strategic Reasoning in Cybersecurity through Large Language Models | Paper Link
- LoRA: Low-Rank Adaptation of Large Language Models | Paper Link
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs | Paper Link
- AI Cyber Risk Benchmark: Automated Exploitation Capabilities | Paper Link
- LLM Agents can Autonomously Hack Websites | Paper Link
- When LLMs Go Online: The Emerging Threat of Web-Enabled LLMs | Paper Link
- Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models | Paper Link
- CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models | Paper Link
B. Multi-agent Collaboration
- BreachSeek: A Multi-Agent Automated Penetration Tester | Paper Link
- PENTEST-AI, an LLM-Powered Multi-Agents Framework for Penetration Testing Automation Leveraging Mitre Attack | Paper Link
- VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework | Paper Link
- Audit-LLM: Multi-Agent Collaboration for Log-based Insider Threat Detection | Paper Link
- Reinforcement Learning-Driven LLM Agent for Automated Attacks on LLMs | Paper Link
III. Common Cyberattacks and Benchmarks of LLM-based Agents
A1. Cyber Threat Intelligence
- Using LLMs to Automate Threat Intelligence Analysis Workflows in Security Operation Centers | Paper Link
- Actionable Cyber Threat Intelligence using Knowledge Graphs and Large Language Modelss | Paper Link
- Cyber Knowledge Completion Using Large Language Models | Paper Link
- Exploring RAG-based Vulnerability Augmentation with LLMs | Paper Link
- LOCALINTEL: Generating Organizational Threat Intelligence from Global and Local Cyber Knowledge | Paper Link
- The Use of Large Language Models (LLM) for Cyber Threat Intelligence (CTI) in Cybercrime Forums | Paper Link
- CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence | Paper Link
A2. Penetration Testing
- Construction and Evaluation of LLM-based agents for Semi-Autonomous penetration testing | Paper Link
- PentestAgent: Incorporating LLM Agents to Automated Penetration Testing | Paper Link
- ARACNE: An LLM-Based Autonomous Shell Pentesting Agent | Paper Link
- Hacking, The Lazy Way: LLM Augmented Pentesting | Paper Link
- AutoPT: How Far Are We from the End2End Automated Web Penetration Testing? | Paper Link
- CIPHER: Cybersecurity Intelligent Penetration-testing Helper for Ethical Researcher | Paper Link
- PenTest++: Elevating Ethical Hacking with AI and Automation | Paper Link
- Getting pwn'd by AI: Penetration Testing with Large Language Models | Paper Link
- PentestGPT: Evaluating and Harnessing Large Language Models for Automated Penetration Testing | Paper Link
- Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks | Paper Link
- RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents | Paper Link
- Penhealnet: An Agent-Based Llm Framework for Automated Pentesting and Optimal Remediation | Paper Link
- PenHeal: A Two-Stage LLM Framework for Automated Pentesting and Optimal Remediation | Paper Link
- BreachSeek: A Multi-Agent Automated Penetration Tester | Paper Link
- PENTEST-AI, an LLM-Powered Multi-Agents Framework for Penetration Testing Automation Leveraging Mitre Attack | Paper Link
- VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework | Paper Link
- A Unified Modeling Framework for Automated Penetration Testing | Paper Link
- AutoPenBench: Benchmarking Generative Agents for Penetration Testing | Paper Link
- Towards Automated Penetration Testing: Introducing LLM Benchmark, Analysis, and Improvements | Paper Link
- HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing | Paper Link
A3. Vulnerability Detection
- LProtector: An LLM-driven Vulnerability Detection System | Paper Link
- WitheredLeaf: Finding Entity-Inconsistency Bugs with LLMs | Paper Link
- Assessing Cybersecurity Vulnerabilities in Code Large Language Models | Paper Link
- Vulnerability Detection and Monitoring Using LLM | Paper Link
- Exploring RAG-based Vulnerability Augmentation with LLMs | Paper Link
- GRACE: Empowering LLM-based software vulnerability detection with graph structure and in-context learning | Paper Link
- Guessing as a service: large language models are not yet ready for vulnerability detection | Paper Link
A4. Phishing and Social Engineering
- Assessing AI vs Human-Authored Spear Phishing SMS Attacks: An Empirical Study | Paper Link
- Cyberattacks Using ChatGPT: Exploring Malicious Content Generation Through Prompt Engineering | Paper Link
- Exploring the Dark Side of AI: Advanced Phishing Attack Design and Deployment Using ChatGPT | Paper Link
- From Chatbots to PhishBots? -- Preventing Phishing scams created using ChatGPT, Google Bard and Claude | Paper Link
- PEEK: Phishing Evolution Framework for Phishing Generation and Evolving Pattern Analysis using Large Language Models | Paper Link
- Next-Generation Phishing: How LLM Agents Empower Cyber Attackers | Paper Link
- On the Feasibility of Fully AI-automated Vishing Attacks | Paper Link
- PhishAgent: A Robust Multimodal Agent for Phishing Webpage Detection | Paper Link
- Anatomy of an AI-powered malicious social botnet | Paper Link
- The Shadow of Fraud: The Emerging Danger of AI-powered Social Engineering and its Possible Cure | Paper Link
- Defending Against Social Engineering Attacks in the Age of LLMs | Paper Link
- Personalized Attacks of Social Engineering in Multi-turn Conversations -- LLM Agents for Simulation and Detection | Paper Link
B1. Malware Generation
- From LLMs to LLM-based Agents for Software Engineering: A Survey of Current, Challenges and Future | Paper Link
- LLM-Based Multi-Agent Systems for Software Engineering: Literature Review, Vision and the Road Ahead | Paper Link
- Wormgpt: a large language model chatbot for criminals | Paper Link
- From Text to MITRE Techniques: Exploring the Malicious Use of Large Language Models for Generating Cyber Attack Payloads | Paper Link
- Tactics, Techniques, and Procedures (TTPs) in Interpreted Malware: A Zero-Shot Generation with Large Language Models | Paper Link
- Assessing LLMs in malicious code deobfuscation of real-world malware campaigns | Paper Link
- Malla: Demystifying Real-world Large Language Model Integrated Malicious Services | Paper Link
- RatGPT: Turning online LLMs into Proxies for Malware Attacks | Paper Link
- AppPoet: Large Language Model based Android malware detection via multi-view prompt engineering | Paper Link
- RedCode: Risky Code Execution and Generation Benchmark for Code Agents | Paper Link
- RedCodeAgent: Automatic Red-teaming Agent against Code Agents | Paper Link
B2. Vulnerability Exploitation
- Leveraging LLM for Zero-Day Exploit Detection in Cloud Networks | Paper Link
- LLM Agents can Autonomously Exploit One-day Vulnerabilities | Paper Link
- Prompting Is All You Need: Automated Android Bug Replay with Large Language Models | Paper Link
- SecureFalcon: Are We There Yet in Automated Software Vulnerability Detection with LLMs? | Paper Link
- Finetuning Large Language Models for Vulnerability Detection | Paper Link
- Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG | Paper Link
- CVE-LLM : Automatic vulnerability evaluation in medical device industry using large language models | Paper Link
B3. Honeypot Deployment
- Act as a Honeytoken Generator! An Investigation into Honeytoken Generation with Large Language Models | Paper Link
- HoneyLLM: A Large Language Model-Powered Medium-Interaction Honeypot | Paper Link
- LLM Honeypot: Leveraging Large Language Models as Advanced Interactive Honeypot Systems | Paper Link
- LLM in the Shell: Generative Honeypots | Paper Link
- LLMPot: Dynamically Configured LLM-based Honeypot for Industrial Protocol and Physical Process Emulation | Paper Link
- LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild | Paper Link
B4. Capture the Flag Challenges
- Using Large Language Models for Cybersecurity Capture-The-Flag Challenges and Certification Questions | Paper Link
- Hacking CTFs with Plain Agents | Paper Link
- Language Agents as Hackers: Evaluating Cybersecurity Skills with Capture the Flag | Paper Link
- EnIGMA: Enhanced Interactive Generative Model Agent for CTF Challenges | Paper Link
IV. Cyberattack Capabilities of LLM-based Agents on Static Infrastructure Networks
A. 6G Core & Radio Access Networks
- Enhancing Network Management Using Code Generated by Large Language Models | Paper Link
- Large language models in 6G security: challenges and opportunities | Paper Link
- On the Feasibility of Using LLMs to Execute Multistage Network Attacks | Paper Link
- Enhancing Autonomous System Security and Resilience With Generative AI: A Comprehensive Survey | Paper Link
- Critical Infrastructure Protection: Generative AI, Challenges, and Opportunities | Paper Link
- Large Language Models to Enhance Malware Detection in Edge Computing | Paper Link
- Large Language Models in Wireless Application Design: In-Context Learning-enhanced Automatic Network Intrusion Detection | Paper Link
- Investigating cybersecurity incidents using large language models in latest-generation wireless networks | Paper Link
B. Enterprise Networks
- A Survey on Enterprise Network Security: Asset Behavioral Monitoring and Distributed Attack Detection | Paper Link
- Can LLMs Hack Enterprise Networks? Autonomous Assumed Breach Penetration-Testing Active Directory Networks | Paper Link
- Leveraging LLM for Zero-Day Exploit Detection in Cloud Networks | Paper Link
D. Software-Defined Networking
- A Systematic Literature Review on Cyber Attack Detection in Software-Define Networking (SDN) | Paper Link
- Identifying cyber-attacks on software defined networks: An inference-based intrusion detection approach | Paper Link
- Cyberattack impact reduction using software-defined networking for cyber-physical production systems | Paper Link
- Unseen Attack Detection in Software-Defined Networking Using a BERT-Based Large Language Model | Paper Link
E. Smart Grids
- Securing smart grid: cyber attacks, countermeasures, and challenges | Paper Link
- A Comprehensive Review on Cyber-Attacks in Power Systems: Impact Analysis, Detection, and Cyber Security | Paper Link
- GridAttackSim: A Cyber Attack Simulation Framework for Smart Grids | Paper Link
- GridAttackAnalyzer: A Cyber Attack Analysis Framework for Smart Grids | Paper Link
- Online Cyber-Attack Detection in Smart Grid: A Reinforcement Learning Approach | Paper Link
- ChatGPT and Other Large Language Models for Cybersecurity of Smart Grid Applications | Paper Link
- Exploring the emerging role of large language models in smart grid cybersecurity: a survey of attacks, detection mechanisms, and mitigation strategies | Paper Link
- Risks of Practicing Large Language Models in Smart Grid: Threat Modeling and Validation | Paper Link
- Vulnerability of Machine Learning Approaches Applied in IoT-based Smart Grid: A Review | Paper Link
F. Quantum Networks
- Applications of LLMs in Quantum-Aware Cybersecurity Leveraging LLMs for Real-Time Anomaly Detection and Threat Intelligence | Paper Link
V. Cyberattack Capabilities of LLM-based Agents on Mobile Infrastructure Networks
A. Internet of Things
- Adversarial Attacks on IoT Systems Leveraging Large Language Models | Paper Link
- ChatIoT: Large Language Model-based Security Assistant for Internet of Things with Retrieval-Augmented Generation | Paper Link
- BARTPredict: Empowering IoT Security with LLM-Driven Cyber Threat Prediction | Paper Link
- Revolutionizing Cyber Threat Detection with Large Language Models: A privacy-preserving BERT-based Lightweight Model for IoT/IIoT Devices | Paper Link
- AttackLLM: LLM-based Attack Pattern Generation for an Industrial Control System | Paper Link
- IoT Vulnerability Detection using Featureless LLM CyBert Model | Paper Link
- LLMPot: Dynamically Configured LLM-based Honeypot for Industrial Protocol and Physical Process Emulation | Paper Link
- Beyond Detection: Leveraging Large Language Models for Cyber Attack Prediction in IoT Networks | Paper Link
B. Satellite Networks
- PLLM-CS: Pre-trained Large Language Model (LLM) for Cyber Threat Detection in Satellite Networks | Paper Link
- Detection of Zero-Day Attacks in a Software-Defined LEO Constellation Network Using Enhanced Network Metric Predictions | Paper Link
C. Mobile Ad-Hoc Networks
- Detection and Evaluation of Cybersecurity Threats in MANET Based on AI | Paper Link
- Using Artificial Intelligence to Evaluating Detection of Cybersecurity Threats in Ad Hoc Networks | Paper Link
- Generative AI-Enhanced Intrusion Detection Framework for Secure Healthcare Networks in MANETs | Paper Link
D. Vehicle Networks
- GenAI-Driven Cyberattack Detection in V2X Networks for Enhanced Road Safety and Autonomous Vehicle Defense | Paper Link
- AI-Based Sensor Attack Detection and Classification for Autonomous Vehicles in 6G-V2X Environment | Paper Link
- Enhancing In-Vehicle Network Security Against AI-Generated Cyberattacks Using Machine Learning | Paper Link
- Artificial Intelligence techniques to mitigate cyber-attacks within vehicular networks: Survey | Paper Link
- AI-Based Intrusion Detection Systems for In-Vehicle Networks: A Survey | Paper Link
- Attacks to Automatous Vehicles: A Deep Learning Algorithm for Cybersecurity | Paper Link
- The Dark Side of AI: Large Language Models as Tools for Cyber Attacks on Vehicle Systems | Paper Link
E. UAV Networks
- Net-GPT: A LLM-Empowered Man-in-the-Middle Chatbot for Unmanned Aerial Vehicle | Paper Link
- How to Detect Cyber-Attacks in Unmanned Aerial Vehicles Network? | Paper Link
- A Survey of Cyberattack Countermeasures for Unmanned Aerial Vehicles | Paper Link
- A Hierarchical Detection and Response System to Enhance Security Against Lethal Cyber-Attacks in UAV Networks | Paper Link
- Unmanned Aerial Vehicles: Vulnerability to Cyber Attacks | Paper Link
- Cyber Attacks on Unmanned Aerial Vehicles and Cyber Security Measures | Paper Link
F. Underwater Networks
- Enhanced SVM and RNN Classifier for Cyberattacks Detection in Underwater Wireless Sensor Networks | Paper Link
- Network Security Risks and Solutions Through Automated Toolkits in Underwater Sensor Network: A Survey | Paper Link
- State-of-the-art security schemes for the Internet of Underwater Things: A holistic survey | Paper Link
VI. Cyberattack Capabilities of LLM-based Agents on Infrastructure-Free Networks
A. Social Networks
- Anatomy of an AI-powered malicious social botnet | Paper Link
- CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM Agent | Paper Link
- Creation and management of social network honeypots for detecting targeted cyber attacks | Paper Link
- Stopping the cyberattack in the early stage: assessing the security risks of social network users | Paper Link
B. Content-Delivery Networks
- Infrastructure upgrade framework for Content Delivery Networks robust to targeted attacks | Paper Link
- DDoS Attack Information Sharing Among CDNs Interconnected Through CDNI | Paper Link
- Investigating Impact of DDoS Attack and CPA Targeting CDN Caches | Paper Link
- Models for Cloud System Availability Assessment Considering Attacks on CDN and ML Based Parametrization | Paper Link
C. Blockchain Networks
- Logic Meets Magic: LLMs Cracking Smart Contract Vulnerabilities | Paper Link
- Cybersecurity Attacks and Detection Methods in Web 3.0 Technology: A Review | Paper Link
- Collaborative Learning for Cyberattack Detection in Blockchain Networks| Paper Link
D. Digital Twin Networks
- Smart Grid: Cyber Attacks, Critical Defense Approaches, and Digital Twin | Paper Link
- Digital Twin-Based Cyber-Attack Detection Framework for Cyber-Physical Manufacturing Systems | Paper Link
- Cyber Attacks on Avionics Networks in Digital Twin Environment: Detection and Defense | Paper Link
- CyberDefender: an integrated intelligent defense framework for digital-twin-based industrial cyber-physical systems | Paper Link
E. Immersive Networks
- Revolutionizing QoE-Driven Network Management with Digital Agents in 6G | Paper Link
- Cyber Security Threats and Challenges in Collaborative Mixed-Reality | Paper Link
- Securing the Virtual Realm: Strategies for Cybersecurity in Augmented Reality (AR) and Virtual Reality (VR) Applications | Paper Link
- Detecting and Preventing Faked Mixed Reality | Paper Link
- Effects of Cyberattacks on Virtual Reality and Augmented Reality Technologies for People with Disabilities | Paper Link
- On the Feasibility of Using MultiModal LLMs to Execute AR Social Engineering Attacks | Paper Link
F. Autonomous Agent Networks
- Emerging Security Challenges of Large Language Models | Paper Link
- Threat Modelling and Risk Analysis for Large Language Model (LLM)-Powered Applications | Paper Link
- Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities | Paper Link
- Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks | Paper Link
- Reinforcement Learning-Driven LLM Agent for Automated Attacks on LLMs | Paper Link
Maintainers
Minrui Xu MINRUI001@e.ntu.edu.sg
Jiani Fan JIANI001@e.ntu.edu.sg
Xinyu Huang x357huan@uwaterloo.ca
Citation
If you find this survey useful, please cite our paper:
@article{xu2025forewarnedforearmedsurveylarge,
title={Forewarned is Forearmed: A Survey on Large Language Model-based Agents in Autonomous Cyberattacks},
author={Xu, Minrui and Fan, Jiani and Huang, Xinyu and Zhou, Conghao and Kang, Jiawen and Niyato, Dusit and Mao, Shiwen and Han, Zhu and Shen, Xuemin (Sherman) and Lam, Kwok-Yan},
journal = {arXiv preprint arXiv:2505.12786},
year={2025},
}
How to Contribute
If you have a paper or are aware of relevant research that should be incorporated, please contribute via pull requests, issues, email, or other suitable methods.