Awesome Chaos Engineering 
Testing in production (TiP) is gaining steam as an accepted practice in DevOps and testing communities, but no amount of preproduction QA testing can foresee all the possible scenarios in your real production deployment.
The prevailing wisdom is that you will see failures in production; the only question is whether you'll be surprised by them or inflict them intentionally to test system resilience and learn from the experience.
The latter approach is chaos engineering.
To understand all this knowledge is very important have a good background in Chaos Engineering, containers, fault injection, monitoring and observability.
Contents
0. Introduction
Chaos engineering is defined as "the discipline of experimenting on a system in order to build confidence in the system's capability to withstand turbulent conditions in production" (Principles of Chaos Engineering).
In other words, it's a software testing method focusing on finding evidence of problems before they are experienced by users.
It's a common misconception that chaos engineering is only about randomly breaking things in production. It's not. Although running experiments in production is a unique part of chaos engineering (more on that later), it's about much more than that—anything that helps us be confident the system can withstand turbulence.
IMPORTANT!: Chaos engineering is not just about randomly breaking things ;-)
It interfaces with site reliability engineering (SRE), application and systems performance analysis, and other forms of testing.
Practicing chaos engineering can help you prepare for failure, and by doing that, learn to build better systems, improve existing ones, and make the world a safer place.
Motivations for chaos engineering
There are at least three good reasons to implement chaos engineering:
- Determining risk and cost and setting service-level indicators, objectives, and agreements
- Testing a system (often complex and distributed) as a whole
- Finding emergent properties you were unaware of
1. Chaos in Practice
To specifically address the uncertainty of distributed systems at scale, Chaos Engineering can be thought of as the facilitation of experiments to uncover systemic weaknesses. These experiments follow four steps:
- Start by defining 'steady state' as some measurable output of a system that indicates normal behavior.
- Hypothesize that this steady state will continue in both the control group and the experimental group.
- Introduce variables that reflect real world events like servers that crash, hard drives that malfunction, network connections that are severed, etc.
- Try to disprove the hypothesis by looking for a difference in steady state between the control group and the experimental group.
The harder it is to disrupt the steady state, the more confidence we have in the behavior of the system. If a weakness is uncovered, we now have a target for improvement before that behavior manifests in the system at large.
2. Principles of Chaos Engineering
A chaos experiment is defined as the following five points by the Principles of chaos engineering
- Build a Hypothesis around Steady State Behavior
- Vary Real-world Events
- Run Experiments in Production
- Automate Experiments to Run Continuously
- Minimize Blast Radius
More details in the following link ;-)
3. Fault Injection
- Chaos Monkey - A resiliency tool that helps applications tolerate random instance failures.
- Chaos Toolkit - A chaos engineering toolkit to help you build confidence in your software system.
- Chaos Toolkit Turbulence - Extension for Chaos Toolkit which adds support for Turbulence attacks.
- Monarch - A series of tools for Chaos Toolkit.
- Muxy - A chaos testing tool for simulating real-world distributed system failures.
- Chaos Blade - Experimental tool that follows the principles of Chaos Engineering to simulate common fault scenarios.
- Cthulhu - Chaos Engineering tool that helps evaluating the resiliency of microservice systems in a data-driven manner.
- Namazu - Programmable fuzzy scheduler for testing distributed systems.
- Chaos Scimmia - Chaos Engineering for Redis.
- HavocLeopard - A set of simple chaos engineering apps for on-prem servers.
- AWS Chaos Scripts - Collection of Python scripts to run failure injection on AWS infrastructure.
CPU
- CPU Troll - Dedicated to raising CPU latency by the requested percentage and timespan.
- stress-ng - Tool to load and stress a computer system in various selectable ways, including CPU, memory, I/O, and more.
Memory
- totalChaos - Overloads RAM and simulates system resource exhaustion scenarios.
Disk
- Disk Fillup - LitmusChaos experiment to fill up disk space on a target node to test storage-related failure handling.
Networking
- Toxiproxy - A TCP proxy to simulate network and system conditions for chaos and resiliency testing.
- Comcast - A tool designed to simulate common network problems like latency, bandwidth restrictions, and dropped/reordered/corrupted packets.
- Chaos HTTP Proxy - Introduces failures into HTTP requests via a proxy server.
Security
- Infection Monkey - Open source security tool for testing a data center's resiliency to perimeter breaches and internal server infection.
- ChaoSlingr - Security Chaos Engineering focused on experimentation on AWS Infrastructure.
- Mitigant - Security chaos engineering for cloud cyber resilience.
Languages
Compilation Time
- ChaosCat - Chaos engineering for Pull Requests - Taking a not-even-good joke a bit too far.
Runtime
- Byteman - A Swiss Army Knife for Byte Code Manipulation.
- Byte-Monkey - Bytecode-level fault injection for the JVM via instrumentation.
- Perses - A project to cause controlled destruction to a JVM application.
- Wiremock - API mocking (Service Virtualization) which enables modeling real world faults and delays.
- MockLab - API mocking (Service Virtualization) as a service.
- Flaw - Injects failures on API calls for local chaos engineering.
- Havoc - Collection of dangerous code that wreaks havoc in .NET applications for chaos-engineering.
- Utilities for frontend chaos engineering - Collection of utilities for applying chaos to frontend applications.
- CHAOS GOPHER - A collection of Unix-style tools in Go for chaos engineering or testing.
- Chaos Monkey for Spring Boot - Injects latencies, exceptions, and terminations into Spring Boot applications.
- React Chaos - Chaos Engineering for your React apps.
- Vue Chaos - A simple yet chaotic component to introduce chaos in your Vue app.
- Chaos QoaLa - Chaos engineering tool for injecting failure into JavaScript-backed GraphQL endpoints.
- Chaos Reverse-engineering - Chaos engineering approach by reverse-engineering.
- Fault - Go HTTP middleware that makes it easy to inject faults into your service.
- GORM SQLChaos - Manipulates DML at program runtime based on GORM callbacks.
- Chaos Frontend Toolkit - A set of tools to break your web apps and find ways to improve them.
Database
- RedFI - Acts as a proxy between the client and Redis with the capability of injecting faults on the fly.
Virtual Machine
- ChaosMachine - Tool to do chaos engineering at the application level in the JVM.
- TripleAgent - System for fault injection for Java applications.
Containers & Orchestrators
- ChaosOrca - Tool for doing Chaos Engineering on containers by perturbing system calls.
- POBS - Automatic Observability and Chaos for Dockerized Java Applications.
- Pumba - Chaos testing and network emulation for Docker containers and clusters.
- Blockade - Docker-based utility for testing network failures and partitions in distributed applications.
- Chaos Engineering for Docker - Chaos engineering experiments for Docker environments.
- Chaos Engineering with Docker EE - Chaos engineering experiments targeting Docker Enterprise Edition.
- Chaos Util - Docker image with utilities for Chaos Engineering.
- Drax - DC/OS Resilience Automated Xenodiagnosis tool for testing DC/OS deployments.
- Pod-Reaper - A rules-based pod killing container for Chaos testing in Kubernetes.
- Chaoskube - Periodically kills random pods in your Kubernetes cluster.
- Litmus - Framework for Kubernetes environments that enables users to run test suites, capture logs, generate reports and perform chaos tests.
- Chaos Operator - Chaos engineering via Kubernetes operator.
- Kube Entropy - A little chaos engineering application for Kubernetes resilience testing.
- kubernetes-chaos-lab - A brief guide to setting up your first chaos engineering lab on Kubernetes.
- Chaos Mesh - A Chaos Engineering Platform for Kubernetes.
Hypervisors
- Turbulence - Tool focused on BOSH environments capable of stressing VMs and manipulating network traffic.
- Chaos Lemur - Self-hostable application to randomly destroy virtual machines in a BOSH-managed environment.
Kernel & Operating System
- stress-ng - Stress test tool for Linux systems covering CPU, memory, I/O, network, and kernel-level stressors.
- sysdig - Linux system exploration and troubleshooting tool with first-class support for containers.
Cloud
- Chaos Engine - Application for creating random Chaos Events in cloud applications to test resiliency.
Private Cloud
- Glooshot - Chaos engineering framework to help you immunize your service mesh.
- kube-monkey - An implementation of Netflix's Chaos Monkey for Kubernetes clusters.
- Powerful Seal - Adds chaos to your Kubernetes clusters, killing targeted pods and taking VMs up and down.
- KubeInvaders - Gamified chaos engineering tool for Kubernetes clusters.
- Kube DOOM - Kill pods inside your Kubernetes cluster by shooting them in Doom.
- GomJabbar - ChaosMonkey for your private cloud.
- kubethanos - Kills half of your pods randomly to engineer chaos in your preferred environment.
- krkn - Chaos and resiliency testing tool for Kubernetes and OpenShift.
- kube-burner - Kubernetes performance and scale test orchestration toolset.
- Chaos Controller - Kubernetes controller for injecting various systemic failures at scale.
Amazon AWS
Azure Cloud
Example Projects
4. Observability
6. Cost of SEVs
7. Chaos as a Service
- Gremlin Inc. - Failure as a Service.
- Chaos Engineering Experiment Automation - Automating and orchestrating chaos engineering experiments.
- Pystol - The cloud chaos engineering toolbox, open source fault injection platform.
- Chaos Platform - Chaos Engineering Platform for Everyone.
- steadybit - Chaos Engineering platform that helps to proactively reduce downtime and provide visibility into systems.
- Cavisson - Chaos engineering platform for resilience testing.
8. Gamedays
9. Books
10. Conferences and Talks
11. Forums and Groups
12. References
13. License

14. Contributing
Contributions welcome! Read the contribution guidelines first.
Thank you!