joshkumar50/hindsight-incident-agent

Python

0

0 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built an autonomous Kubernetes SRE agent that recalls past incidents using Hindsight memory (r/SideProject)

I've been building an autonomous SRE agent for Kubernetes that learns from past incidents using Hindsight persistent memory. Wrote up how the recall/retain loop works and why memory beats naive LLM reasoning for infrastructure incidents. Full write-up:…

1

Sep 29, 2026

README

⚡ Hindsight Incident Agent — Autonomous Kubernetes SRE & Root Cause Intelligence Platform

AI-powered Autonomous SRE Platform for Kubernetes Clusters. Automatically diagnoses, remediates, and learns from incidents using advanced telemetry, chaos engineering, and LLM reasoning. Hindsight Incident Agent integrates Vectorize Hindsight persistent memory to learn from every resolved incident.


🎬 Demonstration

Hindsight Incident Agent Demo


📢 Published Content

Demo Video

Technical Articles

LinkedIn Posts


🏗️ Architecture

Hindsight Incident Agent Architecture

The platform runs 17 microservices on Kubernetes, decoupled through Redis Streams. The control plane is isolated from the application plane with NetworkPolicies. Only the recovery engine has permission to mutate Kubernetes resources, enforced through least-privilege RBAC. Telemetry flows through OpenTelemetry to Prometheus, Loki, and Jaeger.


🌍 Domain & Problem Statement

The Real Business Problem: When production is down, every minute of downtime costs enterprises thousands of dollars. Kubernetes clusters are highly complex, and SRE teams must scramble to manually correlate metrics, logs, and traces across disjointed dashboards to find the root cause. The biggest bottleneck? Stateless troubleshooting. Once an incident is resolved, the knowledge of how it was fixed is often lost in chat logs. When the same issue strikes a month later, engineers waste hours reinventing the wheel because traditional diagnostic tools have no memory of past outages.

The Solution (Hindsight Incident Agent): Hindsight Incident Agent is a Level-3 Autonomous AI Incident Response Agent that makes persistent memory the star. Using a complex ReAct (Reasoning & Acting) loop, it autonomously queries telemetry systems (Prometheus, Elasticsearch) to diagnose root causes. It integrates Vectorize Hindsight — Persistent Agent Memory to learn from past incidents. When an outage is resolved, Hindsight Incident Agent saves the diagnostic fingerprint and playbook. If a semantically similar issue happens again, Hindsight Incident Agent recalls the exact resolution instantly—dropping MTTR (Mean Time To Resolution) drastically.


📂 Project Structure

The platform is organized into 5 primary pillars, avoiding monolith structures in favor of modular, event-driven microservices:

Hindsight-Incident-Agent/
│
├── ui/                 # ⚛️ React 19 Frontend (Vite + shadcn/ui)
│   ├── src/
│   │   ├── pages/      # Dashboard, Incident Center, AI Analysis, Memory Bank
│   │   ├── api/        # Axios API clients connecting to dashboard-bff
│   │   └── components/ # Reusable UI components
│
├── platform/           # 🧠 17 SRE Microservices (FastAPI)
│   ├── dashboard-bff/  # Backend-For-Frontend (routes UI traffic)
│   ├── ai-copilot/     # AI LLM reasoning engine for root-cause explanations
│   ├── chaos-engine/   # Automated chaos injection and stress testing
│   ├── incident-engine/# Manages active alerts and resolutions
│   └── ...             # 13 other specialized engines (monitoring, execution, rollback)
│
├── apps/               # 🎯 Target Dummy Applications (to inject chaos into)
│   ├── auth-service/
│   ├── payment-service/
│   ├── order-service/
│   ├── inventory-service/
│   ├── notification-service/
│   └── traffic-generator/
│
├── infra/              # 🏗️ Infrastructure as Code
│   ├── helm/           # Helm charts for deploying Hindsight Incident Agent components
│   └── manifests/      # Kubernetes manifests (Deployments, Services)
│
├── docs/               # 📖 Architecture and operations guides
├── scripts/            # 🛠️ Build scripts and base Dockerfiles
└── pkg/                # 📦 Shared libraries (Core, EventBus, Telemetry, Math)

🔗 How the Microservices Connect

The architecture is highly decoupled, relying heavily on a Redis EventBus and a Backend-For-Frontend (BFF) pattern.

┌──────────────────┐       GET /api/*        ┌─────────────────────────┐
│                  │ ──────────────────────▶ │                         │
│  React UI (Vite) │                         │  dashboard-bff (Port 80)│
│  (Port 5173)     │ ◀────────────────────── │  (K8s Service)          │
│                  │       JSON Response     │                         │
└──────────────────┘                         └────────┬────────┬───────┘
                                                      │        │
                     ┌────────────────────────────────┘        └────────────────┐
                     ▼                                                  ▼
           ┌──────────────────┐                               ┌────────────────────────────────┐
           │ incident-engine  │                               │ Hindsight Cloud (Vectorize.io) │
           │ (Active Incidents)│                              │ (Root Cause LLM)               │
           └──────────────────┘                               └────────────────────────────────┘

Connection Points in Code:

FromToFileCode
UI → BFFGET /api/memory/bankui/src/api/client.tsapiClient.get('/memory/bank')
Vite Dev Proxylocalhost:3001ui/vite.config.tsproxy: { '/api': { target: 'http://localhost:3001' } }
dashboard-bffai-copilotplatform/dashboard-bff/main.pyclient.post("http://ai-copilot.../explain")
EventBus (Redis)Platform Servicespkg/eventbus/client.pyawait bus.publish("incident.detected", data)

🚀 Quick Start (Minikube & Vite)

You do NOT need to build the Docker image and roll out the Kubernetes deployment for every UI change! The Vite dev server automatically hot-reloads your changes.

1. Start Kubernetes Environment

minikube start
kubectl apply -k infra/k8s

2. Run the UI locally

cd ui
npm install
npm run dev
# Frontend runs at http://localhost:5173

3. Deploy Platform Changes

If you modify a python file in platform/:

eval $(minikube docker-env)
docker build -t hindsight-agent/dashboard-bff:latest --build-arg SERVICE_NAME=dashboard-bff --build-arg dir=platform/dashboard-bff -f platform/dashboard-bff/Dockerfile .
kubectl rollout restart deployment/dashboard-bff -n incident-agent-system

🧠 Hindsight Memory Integration

Hindsight memory is the core judging criterion (25% of score) and is implemented end-to-end in real Python code — not mocked.

How it works

StepWhereWhat happens
recall()platform/ai-orchestrator/main.py — coordinate_ai_workflow()Called before the RCA pipeline. If the Hindsight bank returns a match with score >= 0.8, the platform short-circuits: it publishes RECOVERY_PLAN_READY immediately with memory_hit=True and skips the expensive LLM chain entirely (~1.5 s vs ~12 s).
retain()platform/ai-orchestrator/main.py — after RECOVERY_PLAN_READY (miss path)Called after a fresh recovery is published. Commits the new diagnostic fingerprint + playbook to the bank so future similar incidents get a cache hit.
memory_hit flagAll RECOVERY_PLAN_READY event payloadsSet to True on a Hindsight cache hit, False on a miss. Downstream services and the UI use this flag to annotate incidents with their memory-hit status.

Environment variables

HINDSIGHT_API_KEY=<your-key>          # Required — disables memory if missing (logged, not crashed)
HINDSIGHT_BASE_URL=https://api.hindsight.vectorize.io  # Optional, defaults to this value

Knowledge Engine endpoints

EndpointMethodPurpose
/searchPOSTProxies recall() — returns {"historical_matches": []} on miss, never fake data
/retainPOSTProxies retain() — stores resolved incident playbooks
/memory/statsGETReturns {"bank_id": ..., "total_memories": -1} (SDK has no count method)

🧠 Vectorize Hindsight — Persistent Agent Memory

Hindsight Incident Agent's autonomous memory bank ensures that it learns from every incident.

  • Recall Phase: Before analyzing logs, the platform queries the Memory Bank for past similar outages to instantly suggest proven runbooks.
  • Retain Phase: When an incident is solved or human feedback is given, the platform commits the learning to the bank.

This gives Hindsight Incident Agent the experience of a senior SRE, drastically reducing MTTR for recurring infrastructure patterns.

joshkumar50/hindsight-incident-agent

Python

0

0 commits

updated Sep 29, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built an autonomous Kubernetes SRE agent that recalls past incidents using Hindsight memory (r/SideProject)

I've been building an autonomous SRE agent for Kubernetes that learns from past incidents using Hindsight persistent memory. Wrote up how the recall/retain loop works and why memory beats naive LLM reasoning for infrastructure incidents. Full write-up:…

1

Sep 29, 2026

README

⚡ Hindsight Incident Agent — Autonomous Kubernetes SRE & Root Cause Intelligence Platform

AI-powered Autonomous SRE Platform for Kubernetes Clusters. Automatically diagnoses, remediates, and learns from incidents using advanced telemetry, chaos engineering, and LLM reasoning. Hindsight Incident Agent integrates Vectorize Hindsight persistent memory to learn from every resolved incident.


🎬 Demonstration

Hindsight Incident Agent Demo


📢 Published Content

Demo Video

Technical Articles

LinkedIn Posts


🏗️ Architecture

Hindsight Incident Agent Architecture

The platform runs 17 microservices on Kubernetes, decoupled through Redis Streams. The control plane is isolated from the application plane with NetworkPolicies. Only the recovery engine has permission to mutate Kubernetes resources, enforced through least-privilege RBAC. Telemetry flows through OpenTelemetry to Prometheus, Loki, and Jaeger.


🌍 Domain & Problem Statement

The Real Business Problem: When production is down, every minute of downtime costs enterprises thousands of dollars. Kubernetes clusters are highly complex, and SRE teams must scramble to manually correlate metrics, logs, and traces across disjointed dashboards to find the root cause. The biggest bottleneck? Stateless troubleshooting. Once an incident is resolved, the knowledge of how it was fixed is often lost in chat logs. When the same issue strikes a month later, engineers waste hours reinventing the wheel because traditional diagnostic tools have no memory of past outages.

The Solution (Hindsight Incident Agent): Hindsight Incident Agent is a Level-3 Autonomous AI Incident Response Agent that makes persistent memory the star. Using a complex ReAct (Reasoning & Acting) loop, it autonomously queries telemetry systems (Prometheus, Elasticsearch) to diagnose root causes. It integrates Vectorize Hindsight — Persistent Agent Memory to learn from past incidents. When an outage is resolved, Hindsight Incident Agent saves the diagnostic fingerprint and playbook. If a semantically similar issue happens again, Hindsight Incident Agent recalls the exact resolution instantly—dropping MTTR (Mean Time To Resolution) drastically.


📂 Project Structure

The platform is organized into 5 primary pillars, avoiding monolith structures in favor of modular, event-driven microservices:

Hindsight-Incident-Agent/
│
├── ui/                 # ⚛️ React 19 Frontend (Vite + shadcn/ui)
│   ├── src/
│   │   ├── pages/      # Dashboard, Incident Center, AI Analysis, Memory Bank
│   │   ├── api/        # Axios API clients connecting to dashboard-bff
│   │   └── components/ # Reusable UI components
│
├── platform/           # 🧠 17 SRE Microservices (FastAPI)
│   ├── dashboard-bff/  # Backend-For-Frontend (routes UI traffic)
│   ├── ai-copilot/     # AI LLM reasoning engine for root-cause explanations
│   ├── chaos-engine/   # Automated chaos injection and stress testing
│   ├── incident-engine/# Manages active alerts and resolutions
│   └── ...             # 13 other specialized engines (monitoring, execution, rollback)
│
├── apps/               # 🎯 Target Dummy Applications (to inject chaos into)
│   ├── auth-service/
│   ├── payment-service/
│   ├── order-service/
│   ├── inventory-service/
│   ├── notification-service/
│   └── traffic-generator/
│
├── infra/              # 🏗️ Infrastructure as Code
│   ├── helm/           # Helm charts for deploying Hindsight Incident Agent components
│   └── manifests/      # Kubernetes manifests (Deployments, Services)
│
├── docs/               # 📖 Architecture and operations guides
├── scripts/            # 🛠️ Build scripts and base Dockerfiles
└── pkg/                # 📦 Shared libraries (Core, EventBus, Telemetry, Math)

🔗 How the Microservices Connect

The architecture is highly decoupled, relying heavily on a Redis EventBus and a Backend-For-Frontend (BFF) pattern.

┌──────────────────┐       GET /api/*        ┌─────────────────────────┐
│                  │ ──────────────────────▶ │                         │
│  React UI (Vite) │                         │  dashboard-bff (Port 80)│
│  (Port 5173)     │ ◀────────────────────── │  (K8s Service)          │
│                  │       JSON Response     │                         │
└──────────────────┘                         └────────┬────────┬───────┘
                                                      │        │
                     ┌────────────────────────────────┘        └────────────────┐
                     ▼                                                  ▼
           ┌──────────────────┐                               ┌────────────────────────────────┐
           │ incident-engine  │                               │ Hindsight Cloud (Vectorize.io) │
           │ (Active Incidents)│                              │ (Root Cause LLM)               │
           └──────────────────┘                               └────────────────────────────────┘

Connection Points in Code:

FromToFileCode
UI → BFFGET /api/memory/bankui/src/api/client.tsapiClient.get('/memory/bank')
Vite Dev Proxylocalhost:3001ui/vite.config.tsproxy: { '/api': { target: 'http://localhost:3001' } }
dashboard-bffai-copilotplatform/dashboard-bff/main.pyclient.post("http://ai-copilot.../explain")
EventBus (Redis)Platform Servicespkg/eventbus/client.pyawait bus.publish("incident.detected", data)

🚀 Quick Start (Minikube & Vite)

You do NOT need to build the Docker image and roll out the Kubernetes deployment for every UI change! The Vite dev server automatically hot-reloads your changes.

1. Start Kubernetes Environment

minikube start
kubectl apply -k infra/k8s

2. Run the UI locally

cd ui
npm install
npm run dev
# Frontend runs at http://localhost:5173

3. Deploy Platform Changes

If you modify a python file in platform/:

eval $(minikube docker-env)
docker build -t hindsight-agent/dashboard-bff:latest --build-arg SERVICE_NAME=dashboard-bff --build-arg dir=platform/dashboard-bff -f platform/dashboard-bff/Dockerfile .
kubectl rollout restart deployment/dashboard-bff -n incident-agent-system

🧠 Hindsight Memory Integration

Hindsight memory is the core judging criterion (25% of score) and is implemented end-to-end in real Python code — not mocked.

How it works

StepWhereWhat happens
recall()platform/ai-orchestrator/main.py — coordinate_ai_workflow()Called before the RCA pipeline. If the Hindsight bank returns a match with score >= 0.8, the platform short-circuits: it publishes RECOVERY_PLAN_READY immediately with memory_hit=True and skips the expensive LLM chain entirely (~1.5 s vs ~12 s).
retain()platform/ai-orchestrator/main.py — after RECOVERY_PLAN_READY (miss path)Called after a fresh recovery is published. Commits the new diagnostic fingerprint + playbook to the bank so future similar incidents get a cache hit.
memory_hit flagAll RECOVERY_PLAN_READY event payloadsSet to True on a Hindsight cache hit, False on a miss. Downstream services and the UI use this flag to annotate incidents with their memory-hit status.

Environment variables

HINDSIGHT_API_KEY=<your-key>          # Required — disables memory if missing (logged, not crashed)
HINDSIGHT_BASE_URL=https://api.hindsight.vectorize.io  # Optional, defaults to this value

Knowledge Engine endpoints

EndpointMethodPurpose
/searchPOSTProxies recall() — returns {"historical_matches": []} on miss, never fake data
/retainPOSTProxies retain() — stores resolved incident playbooks
/memory/statsGETReturns {"bank_id": ..., "total_memories": -1} (SDK has no count method)

🧠 Vectorize Hindsight — Persistent Agent Memory

Hindsight Incident Agent's autonomous memory bank ensures that it learns from every incident.

  • Recall Phase: Before analyzing logs, the platform queries the Memory Bank for past similar outages to instantly suggest proven runbooks.
  • Retain Phase: When an incident is solved or human feedback is given, the platform commits the learning to the bank.

This gives Hindsight Incident Agent the experience of a senior SRE, drastically reducing MTTR for recurring infrastructure patterns.

Languages

Python

51.0%

TypeScript

35.8%

PowerShell

6.3%

Dockerfile

4.5%

CSS

1.1%