AI-powered Autonomous SRE Platform for Kubernetes Clusters. Automatically diagnoses, remediates, and learns from incidents using advanced telemetry, chaos engineering, and LLM reasoning. Hindsight Incident Agent integrates Vectorize Hindsight persistent memory to learn from every resolved incident.

Demo Video
Technical Articles
LinkedIn Posts

The platform runs 17 microservices on Kubernetes, decoupled through Redis Streams. The control plane is isolated from the application plane with NetworkPolicies. Only the recovery engine has permission to mutate Kubernetes resources, enforced through least-privilege RBAC. Telemetry flows through OpenTelemetry to Prometheus, Loki, and Jaeger.
The Real Business Problem: When production is down, every minute of downtime costs enterprises thousands of dollars. Kubernetes clusters are highly complex, and SRE teams must scramble to manually correlate metrics, logs, and traces across disjointed dashboards to find the root cause. The biggest bottleneck? Stateless troubleshooting. Once an incident is resolved, the knowledge of how it was fixed is often lost in chat logs. When the same issue strikes a month later, engineers waste hours reinventing the wheel because traditional diagnostic tools have no memory of past outages.
The Solution (Hindsight Incident Agent): Hindsight Incident Agent is a Level-3 Autonomous AI Incident Response Agent that makes persistent memory the star. Using a complex ReAct (Reasoning & Acting) loop, it autonomously queries telemetry systems (Prometheus, Elasticsearch) to diagnose root causes. It integrates Vectorize Hindsight — Persistent Agent Memory to learn from past incidents. When an outage is resolved, Hindsight Incident Agent saves the diagnostic fingerprint and playbook. If a semantically similar issue happens again, Hindsight Incident Agent recalls the exact resolution instantly—dropping MTTR (Mean Time To Resolution) drastically.
The platform is organized into 5 primary pillars, avoiding monolith structures in favor of modular, event-driven microservices:
Hindsight-Incident-Agent/
│
├── ui/ # ⚛️ React 19 Frontend (Vite + shadcn/ui)
│ ├── src/
│ │ ├── pages/ # Dashboard, Incident Center, AI Analysis, Memory Bank
│ │ ├── api/ # Axios API clients connecting to dashboard-bff
│ │ └── components/ # Reusable UI components
│
├── platform/ # 🧠 17 SRE Microservices (FastAPI)
│ ├── dashboard-bff/ # Backend-For-Frontend (routes UI traffic)
│ ├── ai-copilot/ # AI LLM reasoning engine for root-cause explanations
│ ├── chaos-engine/ # Automated chaos injection and stress testing
│ ├── incident-engine/# Manages active alerts and resolutions
│ └── ... # 13 other specialized engines (monitoring, execution, rollback)
│
├── apps/ # 🎯 Target Dummy Applications (to inject chaos into)
│ ├── auth-service/
│ ├── payment-service/
│ ├── order-service/
│ ├── inventory-service/
│ ├── notification-service/
│ └── traffic-generator/
│
├── infra/ # 🏗️ Infrastructure as Code
│ ├── helm/ # Helm charts for deploying Hindsight Incident Agent components
│ └── manifests/ # Kubernetes manifests (Deployments, Services)
│
├── docs/ # 📖 Architecture and operations guides
├── scripts/ # 🛠️ Build scripts and base Dockerfiles
└── pkg/ # 📦 Shared libraries (Core, EventBus, Telemetry, Math)
The architecture is highly decoupled, relying heavily on a Redis EventBus and a Backend-For-Frontend (BFF) pattern.
┌──────────────────┐ GET /api/* ┌─────────────────────────┐
│ │ ──────────────────────▶ │ │
│ React UI (Vite) │ │ dashboard-bff (Port 80)│
│ (Port 5173) │ ◀────────────────────── │ (K8s Service) │
│ │ JSON Response │ │
└──────────────────┘ └────────┬────────┬───────┘
│ │
┌────────────────────────────────┘ └────────────────┐
▼ ▼
┌──────────────────┐ ┌────────────────────────────────┐
│ incident-engine │ │ Hindsight Cloud (Vectorize.io) │
│ (Active Incidents)│ │ (Root Cause LLM) │
└──────────────────┘ └────────────────────────────────┘
| From | To | File | Code |
|---|---|---|---|
| UI → BFF | GET /api/memory/bank | ui/src/api/client.ts | apiClient.get('/memory/bank') |
| Vite Dev Proxy | localhost:3001 | ui/vite.config.ts | proxy: { '/api': { target: 'http://localhost:3001' } } |
| dashboard-bff | ai-copilot | platform/dashboard-bff/main.py | client.post("http://ai-copilot.../explain") |
| EventBus (Redis) | Platform Services | pkg/eventbus/client.py | await bus.publish("incident.detected", data) |
You do NOT need to build the Docker image and roll out the Kubernetes deployment for every UI change! The Vite dev server automatically hot-reloads your changes.
minikube start
kubectl apply -k infra/k8s
cd ui
npm install
npm run dev
# Frontend runs at http://localhost:5173
If you modify a python file in platform/:
eval $(minikube docker-env)
docker build -t hindsight-agent/dashboard-bff:latest --build-arg SERVICE_NAME=dashboard-bff --build-arg dir=platform/dashboard-bff -f platform/dashboard-bff/Dockerfile .
kubectl rollout restart deployment/dashboard-bff -n incident-agent-system
Hindsight memory is the core judging criterion (25% of score) and is implemented end-to-end in real Python code — not mocked.
| Step | Where | What happens |
|---|---|---|
| recall() | platform/ai-orchestrator/main.py — coordinate_ai_workflow() | Called before the RCA pipeline. If the Hindsight bank returns a match with score >= 0.8, the platform short-circuits: it publishes RECOVERY_PLAN_READY immediately with memory_hit=True and skips the expensive LLM chain entirely (~1.5 s vs ~12 s). |
| retain() | platform/ai-orchestrator/main.py — after RECOVERY_PLAN_READY (miss path) | Called after a fresh recovery is published. Commits the new diagnostic fingerprint + playbook to the bank so future similar incidents get a cache hit. |
| memory_hit flag | All RECOVERY_PLAN_READY event payloads | Set to True on a Hindsight cache hit, False on a miss. Downstream services and the UI use this flag to annotate incidents with their memory-hit status. |
HINDSIGHT_API_KEY=<your-key> # Required — disables memory if missing (logged, not crashed)
HINDSIGHT_BASE_URL=https://api.hindsight.vectorize.io # Optional, defaults to this value
| Endpoint | Method | Purpose |
|---|---|---|
/search | POST | Proxies recall() — returns {"historical_matches": []} on miss, never fake data |
/retain | POST | Proxies retain() — stores resolved incident playbooks |
/memory/stats | GET | Returns {"bank_id": ..., "total_memories": -1} (SDK has no count method) |
Hindsight Incident Agent's autonomous memory bank ensures that it learns from every incident.
This gives Hindsight Incident Agent the experience of a senior SRE, drastically reducing MTTR for recurring infrastructure patterns.
Python
51.0%
TypeScript
35.8%
PowerShell
6.3%
Dockerfile
4.5%
CSS
1.1%
AI-powered Autonomous SRE Platform for Kubernetes Clusters. Automatically diagnoses, remediates, and learns from incidents using advanced telemetry, chaos engineering, and LLM reasoning. Hindsight Incident Agent integrates Vectorize Hindsight persistent memory to learn from every resolved incident.

Demo Video
Technical Articles
LinkedIn Posts

The platform runs 17 microservices on Kubernetes, decoupled through Redis Streams. The control plane is isolated from the application plane with NetworkPolicies. Only the recovery engine has permission to mutate Kubernetes resources, enforced through least-privilege RBAC. Telemetry flows through OpenTelemetry to Prometheus, Loki, and Jaeger.
The Real Business Problem: When production is down, every minute of downtime costs enterprises thousands of dollars. Kubernetes clusters are highly complex, and SRE teams must scramble to manually correlate metrics, logs, and traces across disjointed dashboards to find the root cause. The biggest bottleneck? Stateless troubleshooting. Once an incident is resolved, the knowledge of how it was fixed is often lost in chat logs. When the same issue strikes a month later, engineers waste hours reinventing the wheel because traditional diagnostic tools have no memory of past outages.
The Solution (Hindsight Incident Agent): Hindsight Incident Agent is a Level-3 Autonomous AI Incident Response Agent that makes persistent memory the star. Using a complex ReAct (Reasoning & Acting) loop, it autonomously queries telemetry systems (Prometheus, Elasticsearch) to diagnose root causes. It integrates Vectorize Hindsight — Persistent Agent Memory to learn from past incidents. When an outage is resolved, Hindsight Incident Agent saves the diagnostic fingerprint and playbook. If a semantically similar issue happens again, Hindsight Incident Agent recalls the exact resolution instantly—dropping MTTR (Mean Time To Resolution) drastically.
The platform is organized into 5 primary pillars, avoiding monolith structures in favor of modular, event-driven microservices:
Hindsight-Incident-Agent/
│
├── ui/ # ⚛️ React 19 Frontend (Vite + shadcn/ui)
│ ├── src/
│ │ ├── pages/ # Dashboard, Incident Center, AI Analysis, Memory Bank
│ │ ├── api/ # Axios API clients connecting to dashboard-bff
│ │ └── components/ # Reusable UI components
│
├── platform/ # 🧠 17 SRE Microservices (FastAPI)
│ ├── dashboard-bff/ # Backend-For-Frontend (routes UI traffic)
│ ├── ai-copilot/ # AI LLM reasoning engine for root-cause explanations
│ ├── chaos-engine/ # Automated chaos injection and stress testing
│ ├── incident-engine/# Manages active alerts and resolutions
│ └── ... # 13 other specialized engines (monitoring, execution, rollback)
│
├── apps/ # 🎯 Target Dummy Applications (to inject chaos into)
│ ├── auth-service/
│ ├── payment-service/
│ ├── order-service/
│ ├── inventory-service/
│ ├── notification-service/
│ └── traffic-generator/
│
├── infra/ # 🏗️ Infrastructure as Code
│ ├── helm/ # Helm charts for deploying Hindsight Incident Agent components
│ └── manifests/ # Kubernetes manifests (Deployments, Services)
│
├── docs/ # 📖 Architecture and operations guides
├── scripts/ # 🛠️ Build scripts and base Dockerfiles
└── pkg/ # 📦 Shared libraries (Core, EventBus, Telemetry, Math)
The architecture is highly decoupled, relying heavily on a Redis EventBus and a Backend-For-Frontend (BFF) pattern.
┌──────────────────┐ GET /api/* ┌─────────────────────────┐
│ │ ──────────────────────▶ │ │
│ React UI (Vite) │ │ dashboard-bff (Port 80)│
│ (Port 5173) │ ◀────────────────────── │ (K8s Service) │
│ │ JSON Response │ │
└──────────────────┘ └────────┬────────┬───────┘
│ │
┌────────────────────────────────┘ └────────────────┐
▼ ▼
┌──────────────────┐ ┌────────────────────────────────┐
│ incident-engine │ │ Hindsight Cloud (Vectorize.io) │
│ (Active Incidents)│ │ (Root Cause LLM) │
└──────────────────┘ └────────────────────────────────┘
| From | To | File | Code |
|---|---|---|---|
| UI → BFF | GET /api/memory/bank | ui/src/api/client.ts | apiClient.get('/memory/bank') |
| Vite Dev Proxy | localhost:3001 | ui/vite.config.ts | proxy: { '/api': { target: 'http://localhost:3001' } } |
| dashboard-bff | ai-copilot | platform/dashboard-bff/main.py | client.post("http://ai-copilot.../explain") |
| EventBus (Redis) | Platform Services | pkg/eventbus/client.py | await bus.publish("incident.detected", data) |
You do NOT need to build the Docker image and roll out the Kubernetes deployment for every UI change! The Vite dev server automatically hot-reloads your changes.
minikube start
kubectl apply -k infra/k8s
cd ui
npm install
npm run dev
# Frontend runs at http://localhost:5173
If you modify a python file in platform/:
eval $(minikube docker-env)
docker build -t hindsight-agent/dashboard-bff:latest --build-arg SERVICE_NAME=dashboard-bff --build-arg dir=platform/dashboard-bff -f platform/dashboard-bff/Dockerfile .
kubectl rollout restart deployment/dashboard-bff -n incident-agent-system
Hindsight memory is the core judging criterion (25% of score) and is implemented end-to-end in real Python code — not mocked.
| Step | Where | What happens |
|---|---|---|
| recall() | platform/ai-orchestrator/main.py — coordinate_ai_workflow() | Called before the RCA pipeline. If the Hindsight bank returns a match with score >= 0.8, the platform short-circuits: it publishes RECOVERY_PLAN_READY immediately with memory_hit=True and skips the expensive LLM chain entirely (~1.5 s vs ~12 s). |
| retain() | platform/ai-orchestrator/main.py — after RECOVERY_PLAN_READY (miss path) | Called after a fresh recovery is published. Commits the new diagnostic fingerprint + playbook to the bank so future similar incidents get a cache hit. |
| memory_hit flag | All RECOVERY_PLAN_READY event payloads | Set to True on a Hindsight cache hit, False on a miss. Downstream services and the UI use this flag to annotate incidents with their memory-hit status. |
HINDSIGHT_API_KEY=<your-key> # Required — disables memory if missing (logged, not crashed)
HINDSIGHT_BASE_URL=https://api.hindsight.vectorize.io # Optional, defaults to this value
| Endpoint | Method | Purpose |
|---|---|---|
/search | POST | Proxies recall() — returns {"historical_matches": []} on miss, never fake data |
/retain | POST | Proxies retain() — stores resolved incident playbooks |
/memory/stats | GET | Returns {"bank_id": ..., "total_memories": -1} (SDK has no count method) |
Hindsight Incident Agent's autonomous memory bank ensures that it learns from every incident.
This gives Hindsight Incident Agent the experience of a senior SRE, drastically reducing MTTR for recurring infrastructure patterns.
Python
51.0%
TypeScript
35.8%
PowerShell
6.3%
Dockerfile
4.5%
CSS
1.1%