tradertanmay/ai-agents-zero-to-hero

Learn AI agents from first principles to production. Zero mandatory dependencies, pure standard Python 3.11+, and no magic frameworks.

Python

10

12 commits

updated Oct 6, 2026

See the code

See what people are saying

README

AI Agents: Zero to Hero

AI Agents: Zero → Hero

A publicly accessible, source-available educational framework by Tanmay Sah.

Everyone is talking about AI agents.

But what actually makes something an agent?

Is ChatGPT an agent? Is a deterministic workflow an agent? What happens between an LLM deciding to call a tool and that tool actually executing? Where does memory live? Who controls the loop? Who decides when the agent stops? And what happens when an agent makes the wrong change?

This repository answers those questions from first principles.

No magic. No framework-first abstractions.


What about AI agents confuses you?

Have a question or a concept you want demystified? Check out QUESTIONS.md or submit a question via GitHub Issues. Questions from the community directly shape upcoming modules in this series!


Where Should You Start?


5-Minute Zero-Dependency Quick Start

Clone and run the complete agent loop immediately with Python 3.11+:

git clone https://github.com/tradertanmay/ai-agents-zero-to-hero.git
cd ai-agents-zero-to-hero

# 1. Compare Chatbot vs Workflow vs Agent
python3 01-what-is-an-agent/example.py

# 2. Run the pure Python Observe-Decide-Act loop
python3 02-agent-loop/example.py

# 3. Run your first complete multi-step agent
python3 04-build-your-first-agent/example.py

No API key. No framework. No dependencies. Just Python.


What This Repository Is

A practical, progressive, code-first curriculum designed to take you from:

"I understand LLMs and APIs, but I don't really understand what people mean by an AI agent."

to:

"I understand how agents work internally and can build, debug, evaluate, and reason about production agent systems."


Who This Is For

This repository is built for everyone who wants to understand and build AI agents — whether you are a software engineer, technical lead, researcher, product builder, student, or curious developer.

It is especially designed for you if you:

  • Want to move beyond prompt engineering and understand how autonomous agent systems actually work
  • Know basic Python (or can follow readable standard code)
  • Have used ChatGPT or called LLM APIs, but want to see the underlying machinery behind tools, memory, and loops
  • Hear industry buzzwords like ReAct, Function Calling, Memory, Agent Harness, Multi-Agent, MCP and want a crystal-clear, framework-independent mental model

Core Philosophy: The Model is Not the Agent

A common beginner assumption is:

$$\text{Agent} \stackrel{?}{=} \text{LLM} + \text{Prompt}$$

In reality, a base LLM inference call does not itself maintain persistent application state across turns. An Agent System is a composite computational system comprising a model policy, an execution control loop, state management, and tools, which observes and acts upon an external Environment.

$$\mathbf{Agent\ System = Model/Policy + Runtime/Control\ Loop + State + Tools}$$

flowchart TD
    subgraph AgentSystem["AGENT SYSTEM"]
        M["Model / Policy (Decision Engine)"]
        R["Runtime / Control Loop (Supervisor)"]
        S["State & Memory"]
        T["Tools & Actions"]
        
        R <--> M
        R <--> S
        R <--> T
    end

    AgentSystem <-->|Act / Observe| E["ENVIRONMENT<br/>(APIs, Filesystem, DB, User)"]

    style AgentSystem fill:#f8f9fa,stroke:#333,stroke-width:2px
    style M fill:#f3e5f5,stroke:#7b1fa2
    style R fill:#fff3e0,stroke:#e65100,stroke-width:2px
    style S fill:#ede7f6,stroke:#4527a0
    style T fill:#e8f5e9,stroke:#2e7d32
    style E fill:#e1f5fe,stroke:#0288d1,stroke-width:2px

The Core Agent Loop

At the heart of every agent is the cyclic feedback loop with its environment:

flowchart LR
    O["1. Observe"] --> D["2. Decide"]
    D --> A["3. Act"]
    A --> O

    style O fill:#e1f5fe,stroke:#0288d1
    style D fill:#f3e5f5,stroke:#7b1fa2
    style A fill:#e8f5e9,stroke:#388e3c
  1. Observe: Read the current state of the world, conversation history, and tool feedback.
  2. Decide: The model reasons over the observation and chooses an action or final response.
  3. Act: An execution layer performs the action on the environment.
  4. Observe Again: The environment output becomes a new observation fed back to the model.

Key Rule: The model chooses or requests an action; an execution layer performs it. In a from-scratch agent like the one in this course, that execution layer is our Python runtime. In hosted platforms, the provider may execute certain hosted tools on the application’s behalf.


What People Call an "Agent" — Operational Spectrum

There is no universally agreed boundary across industry and research for what counts as an "agent." In this course, we use the following operational spectrum to make the architectural differences explicit:

flowchart LR
    LLM["1. Raw LLM"] --> Chat["2. Chatbot"]
    Chat --> RAG["3. RAG Pipeline"]
    RAG --> Work["4. Workflow"]
    Work --> ToolLLM["5. Tool-Using LLM"]
    ToolLLM --> Agent["6. Iterative Autonomous Agent"]
    Agent --> Multi["7. Multi-Agent System"]

    style LLM fill:#f5f5f5,stroke:#9e9e9e
    style Chat fill:#e1f5fe,stroke:#0288d1
    style RAG fill:#e0f7fa,stroke:#0097a7
    style Work fill:#fff8e1,stroke:#f57f17
    style ToolLLM fill:#f3e5f5,stroke:#7b1fa2
    style Agent fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
    style Multi fill:#ede7f6,stroke:#4527a0
StageParadigmHow Decisions Are MadeCan it take actions?Closed Feedback Loop?Agentic characteristics
1Raw LLMPredicts next tokens based on promptNoNoNone
2ChatbotAppends user/assistant turns to historyNoOnly via userLow
3RAG PipelineHardcoded retrieval $\to$ LLM synthesisNoNoLow
4Workflow / ChainFixed code sequence ($A \to B \to C$)Fixed actionsFixed error handlingPredetermined control
5Tool-Using LLMModel selects 1 tool $\to$ returns answerYes 1-shot actionNo iterative retrySome
6Iterative Autonomous AgentDynamic Observe $\to$ Decide $\to$ Act loopYes Dynamic actionsYes Closed runtime feedbackHigh
7Multi-Agent SystemMultiple agents with handoffs & rolesYes Distributed actionsYes Inter-agent feedbackMultiple interacting agents

Conceptual Clarity: Stop Mixing These Up

The 5 Context & Data Concepts

TermWhat It Actually IsLifespan
PromptThe static text template or instruction formatted for a single inference call.Single API call
ContextThe exact collection of tokens passed into the model's attention window at time $t$.Single API call
HistoryThe ordered sequence of prior user/assistant turns and tool observations.Conversation
StateThe full operational data structure (variables, step count, budget, artifacts, scratchpad).Task execution
MemoryKnowledge persisted across tasks/sessions (vector indices, key-value stores, user profiles).Persistent / Multi-session

The 5 System Components

ComponentPrimary ResponsibilityExample
ModelProbabilistic decision-maker and text generator.GPT-4o, Claude 3.5 Sonnet, Gemini 2.0
ToolA callable capability that inspects or mutates the environment.calculator(), read_file(), sql_query()
EnvironmentThe external world the agent observes and acts upon.Filesystem, REST API, Database, Shell
Runtime / HarnessThe supervisor controlling loops, step limits, permissions, and tool execution.Python loop, LangGraph runtime
Agent SystemThe complete composite system (Model + Runtime + Tools + State).Minimal Agent, Coding Assistant

Visual Curriculum Map

Every module is labeled with a difficulty level so you can track your progression:

Level 1: Beginner → Level 2: Builder → Level 3: Systems → Level 4: Production → Level 5: Research
flowchart TD
    subgraph L1["Level 1 — Beginner"]
        M00["00-introduction<br/>(Prerequisites & Setup)"]
        M01["01-what-is-an-agent<br/>(Operational Taxonomy)"]
        M02["02-agent-loop<br/>(Observe-Decide-Act Mechanics)"]
        M00 --> M01 --> M02
    end

    subgraph L2["Level 2 — Builder"]
        M03["03-tools-and-function-calling<br/>(Tool Schemas & MCP)"]
        M04["04-build-your-first-agent<br/>(The Complete End-to-End Agent)"]
        M05["05-state-and-memory<br/>(State, History & Memory Stores)"]
        M02 --> M03 --> M04 --> M05
    end

    subgraph L3["Level 3 — Systems"]
        M06["06-planning-and-reasoning<br/>(ReAct & Task Decomposition)"]
        M07["07-context-engineering<br/>(Window Budgets & Pollution)"]
        M08["08-agent-runtime-and-harness<br/>(Middleware, Budgets & Limits)"]
        M09["09-multi-agent-systems<br/>(Handoffs & Supervisor Patterns)"]
        M05 --> M06 --> M07 --> M08 --> M09
    end

    subgraph L4["Level 4 — Production"]
        M10["10-agent-failures<br/>(Taxonomy & Defense Patterns)"]
        M11["11-agent-evaluation<br/>(Trajectory Evals & Benchmarks)"]
        M12["12-agent-safety-and-verification<br/>(Sandboxing & Approvals)"]
        M13["13-production-agents<br/>(Observability & Persistence)"]
        M09 --> M10 --> M11 --> M12 --> M13
    end

    subgraph L5["Level 5 — Research"]
        M14["14-coding-agents<br/>(Repo Search, Patching & Evals)"]
        M15["15-self-improving-agents<br/>(Meta-Learning & Evolution)"]
        M13 --> M14 --> M15
    end

    style L1 fill:#e8f5e9,stroke:#2e7d32
    style L2 fill:#e1f5fe,stroke:#0288d1
    style L3 fill:#fff8e1,stroke:#f57f17
    style L4 fill:#fbe9e7,stroke:#d84315
    style L5 fill:#f3e5f5,stroke:#6a1b9a

Curriculum Table of Contents (Numerical Progression)

ModuleLevelStatusWhat You Will Learn
00-introductionLevel 1ReadyCourse architecture, prerequisites, mental models
01-what-is-an-agentLevel 1ReadyLLM vs Chatbot vs Workflow vs Agent; operational taxonomy
02-agent-loopLevel 1ReadyThe cyclic Observe-Decide-Act execution loop
03-tools-and-function-callingLevel 2ReadyTool lifecycle, schemas, validation & MCP Deep Dive
04-build-your-first-agentLevel 2ReadyAssembling the first complete agent + 10-case regression scorecard
05-state-and-memoryLevel 2ReadyWorking state, SQLite persistent memory & history compaction
06-planning-and-reasoningLevel 3ReadyReAct, decomposition, reflection, and limits of reasoning
07-context-engineeringLevel 3ReadyContext budgets, compression, and anti-pollution
08-agent-runtime-and-harnessLevel 3ReadyHarness as the OS: step limits, budgets, middleware & aborts
09-multi-agent-systemsLevel 3ReadySupervisor-worker, handoffs, and when NOT to use multi-agent
10-agent-failuresLevel 4ReadyFailure taxonomy, loops, ambiguous writes & reconciliation
11-agent-evaluationLevel 4ReadyAdvanced trajectory evaluation, multi-dimensional metrics & frozen benchmarks
12-agent-safety-and-verificationLevel 4ReadyCapability gating, blast radius, invariants & postconditions
13-production-agentsLevel 4ReadyDurable queues, worker leases, crash recovery, telemetry & health checks
14-coding-agentsLevel 5ReadyRepo search, AST navigation, sandboxed test execution & patch repair loops
15-self-improving-agentsLevel 5ReadyThe 5 adaptation surfaces, isolated candidate evals, multi-dimensional gates & rollback

Zero Dependencies & Framework Independence

We believe you should understand how agents work even if every agent framework disappeared tomorrow.

  • Zero Mandatory Third-Party Packages: Everything runs on standard Python 3.11+.
  • Zero Required API Keys: All core lessons include deterministic simulators and mock LLMs for 100% offline learning and automated testing.
  • Pluggable Real Providers: Want to use a real model? Drop in your API key for OpenAI, Anthropic, Gemini, or local Ollama instances in examples/minimal_agent/.

Once you master the mechanics in this repo, you will understand how modern tools and frameworks organize these responsibilities:

  • LangGraph — maps many of these concepts into a graph-oriented orchestration/runtime model with explicit state, nodes, durable execution, and human-in-the-loop control.
  • AutoGen — provides abstractions for agents, teams, messaging, and event-driven multi-agent orchestration.
  • OpenAI Agents SDK — provides an agent runtime with a built-in agent loop, function tools, handoffs, guardrails, sessions, and tracing.
  • MCP (Model Context Protocol) — an open protocol for connecting AI applications to external capabilities and context providers. MCP servers can expose tools, resources, and prompts through a standardized protocol.

Repository Structure (Strict Numerical Order)

ai-agents-zero-to-hero/
├── README.md # Main course landing page
├── LICENSE # Proprietary License (Tanmay Sah)
├── CONTRIBUTING.md # Contribution & pedagogical standard
├── ROADMAP.md # Curriculum milestone checklist
├── QUESTIONS.md # Community Q&A hub
├── pyproject.toml # Standard Python packaging config
├── .github/ # Issue templates
│
├── 00-introduction/ # Module 0: Prerequisites & mental models
├── 01-what-is-an-agent/ # Module 1: Operational spectrum & agent architecture
├── 02-agent-loop/ # Module 2: The Observe-Decide-Act loop
├── 03-tools-and-function-calling/ # Module 3: Tool schemas, execution lifecycle & MCP
├── 04-build-your-first-agent/ # Module 4: Assembling your first complete agent
├── 05-state-and-memory/ # Module 5: Working State, SQLite Memory & Compaction
├── 06-planning-and-reasoning/ # Module 6: ReAct, Planning & Dynamic Replanning
├── 07-context-engineering/ # Module 7: Token Budgets & Observation Pruning
├── 08-agent-runtime-and-harness/ # Module 8: The Agent Harness / Operating System
├── 09-multi-agent-systems/ # Module 9: Supervisor-Worker & Review Loops
├── 10-agent-failures/             # Module 10: Failure Taxonomy, Ambiguous Writes & Reconciliation
├── 11-agent-evaluation/            # Module 11: Advanced Trajectory Evaluation & Benchmarks
│   ├── README.md                   # Core guide, 5 evaluation levels & commands
│   ├── concepts.md                 # Comprehensive deep-dive & metric formulas
│   ├── eval_cases.json             # Frozen 20-case benchmark test suite
│   ├── example.py                  # Runnable comparative regression benchmark (V1 vs V2)
│   └── exercise.md                 # Production incident reproduction exercise
├── 12-agent-safety-and-verification/ # Module 12: Capability Gating, Blast Radius & Compensation
│   ├── README.md                   # Core guide, 4 layers, 6 controls & commands
│   ├── concepts.md                 # Deep-dive: blast radius formula, 6 invariants & compensation
│   ├── example.py                  # Runnable adversarial safety suite (40+ attack vectors)
│   └── exercise.md                 # Edit comment capability & compensation exercise
├── 13-production-agents/            # Module 13: Durable Queues, Worker Leases & Telemetry
│   ├── README.md                   # Core guide, 6 production concerns & commands
│   ├── concepts.md                 # Deep-dive: durable state machine, reconciliation & health
│   ├── example.py                  # Runnable demo: worker leases, crash recovery & trace waterfalls
│   └── exercise.md                 # Incident root-cause analysis & graceful SIGTERM drain
├── 14-coding-agents/               # Module 14: Repo Search, AST Navigation & Sandboxed Repair
│   ├── README.md                   # Core guide, 9-stage loop, invariants & scorecard
│   ├── concepts.md                 # Deep-dive: AST slicing, execution sandboxes & test immutability
│   ├── example.py                  # Component walkthrough: AST slicing, safety linting & sandboxed repair
│   ├── exercise.md                 # Call-graph extraction, forbidden-pattern security & modulo repair
│   └── demo_repo/                  # Deliberately broken mini-repository (calculator & parser bugs)
├── 15-self-improving-agents/        # Module 15: Self-Improving Agents & Governance
│   ├── README.md                   # Core guide, 5 adaptation surfaces, architecture & commands
│   ├── concepts.md                 # Deep-dive: evaluator isolation, multi-dimensional gates & rollback
│   ├── example.py                  # Walkthrough: holdout boundary, multi-dimensional evals & rollback
│   ├── exercise.md                 # Eval dataset tamper-proofing, AST scanner & canary rollback
│   └── evals/                      # Quarantined eval datasets (development, regression, holdout)
│
├── examples/
│   ├── self_improving_agent/       # Capstone: Versioned Self-Improvement & Rollback Subsystem
│   │   ├── versions.py             # Version registry, promotion statuses & immutable audit trail
│   │   ├── mutation.py             # 5 adaptation surfaces, risk hierarchy & protected verifier guard
│   │   ├── proposer.py             # Failure log analysis & adaptation proposal (holdout isolated)
│   │   ├── candidate.py            # Candidate workspace cloning & isolated patch application
│   │   ├── baseline.py             # Baseline naive agent & candidate AST-localized agent
│   │   ├── evaluator.py            # Frozen regression & holdout benchmark scoring
│   │   ├── promotion.py            # Multi-dimensional gate (zero tolerance on unsafe) & human approval
│   │   ├── rollback.py             # Production rollback manager & ancestor restoration
│   │   └── main.py                 # Flagship end-to-end runnable demonstration
│   ├── coding_assistant/           # Applied Coding Agent: Sandboxed Repair & Verification Subsystem
│   │   ├── repo.py                 # Repository discovery, language detection & tree mapping
│   │   ├── search.py               # Fast code grep & symbol definition lookup
│   │   ├── ast_tools.py            # AST symbol extraction & targeted function slicing
│   │   ├── patch.py                # Unified diffs, targeted replacements & syntax validation
│   │   ├── sandbox.py              # Isolated tempdir cloning, command execution & repo sync
│   │   ├── verifier.py             # Test execution, traceback parsing & safety linter
│   │   ├── agent.py                # Autonomous coding agent loop & verification scorecard
│   │   └── main.py                 # End-to-end runnable demo on demo_repo
│   ├── minimal_agent/              # Modular, runnable showcase agent
│   │   ├── README.md
│   │   ├── llm.py                  # Pluggable LLM interface (Mock, OpenAI, Anthropic, Gemini, Ollama)
│   │   ├── state.py                # State representation & history
│   │   ├── tools.py                # Tool definitions & registry
│   │   ├── runtime.py              # Step controller & budget enforcement
│   │   ├── agent.py                # Pure agent logic
│   │   └── main.py                 # Runnable demo script
│   └── reddit_comment_agent/       # Capstone: Human-in-the-Loop Reddit Comment Agent (Modules 01-09)
│       ├── README.md               # Architecture, module mapping & user guide
│       ├── mock_reddit.py          # Offline simulated Reddit environment
│       ├── reddit.py               # Reddit client interface (Mock + optional PRAW)
│       ├── state.py                # SQLite persistent memory & duplicate prevention
│       ├── tools.py                # Permission-gated tool registry
│       ├── evaluator.py            # 5-criterion quality scorecard
│       ├── approval.py             # Human-in-the-loop review gate & HMAC tokens
│       ├── safety.py               # 4-layer defense, blast limiter, verifiers & compensation
│       ├── runtime.py              # OS harness & rate limits
│       ├── agent.py                # Central coordinator (observe, decide, act)
│       ├── main.py                 # Standalone runnable demo script
│       ├── evals/                  # Benchmark evaluation harness (Module 11)
│       │   ├── cases.json          # Frozen 20-case test suite
│       │   ├── metrics.py          # Multi-dimensional metric calculations
│       │   ├── judges.py           # Deterministic, Heuristic & LLM Judges
│       │   ├── runner.py           # Benchmark execution harness
│       │   └── report.py           # Regression reporting & delta tables
│       └── production/             # Production subsystem (Module 13)
│           ├── config.py           # Configuration, secrets & redaction
│           ├── jobs.py             # State machine & transition validation
│           ├── checkpoints.py      # Durable SQLite queue & worker leases
│           ├── telemetry.py        # Structured JSONL, metrics & tracer
│           ├── health.py           # Liveness, readiness & dependency health
│           ├── recovery.py         # Crash recovery & post-commit reconciliation
│           └── worker.py           # Worker loop & graceful shutdown
└── tests/                          # Unittest verification suite (88 tests)
    ├── test_module_01.py
    ├── test_module_02.py
    ├── test_module_03.py
    ├── test_module_04.py
    ├── test_module_05.py
    ├── test_module_06.py
    ├── test_module_07.py
    ├── test_module_08.py
    ├── test_module_09.py
    ├── test_module_10.py
    ├── test_module_11.py
    ├── test_module_12.py
    ├── test_module_13.py
    ├── test_minimal_agent.py
    └── test_reddit_comment_agent.py

Running the Tests

Run the zero-dependency test suite using standard Python:

python3 -m unittest discover -s tests -v

License

Copyright (c) 2026 Tanmay Sah. All rights reserved.

This work is published under a Proprietary / Source-Available License. No part of the written curriculum, diagrams, code implementations, exercises, or related expressive materials may be reproduced, distributed, modified, or used for commercial, corporate training, or derivative purposes without prior written permission. See LICENSE for full terms.

agent-runtime
ai-agents
autonomous-agents
first-principles
function-calling
llm
mcp
multi-agent
python
zero-to-hero

tradertanmay/ai-agents-zero-to-hero

Learn AI agents from first principles to production. Zero mandatory dependencies, pure standard Python 3.11+, and no magic frameworks.

Python

10

12 commits

updated Oct 6, 2026

See the code

See what people are saying

README

AI Agents: Zero to Hero

AI Agents: Zero → Hero

A publicly accessible, source-available educational framework by Tanmay Sah.

Everyone is talking about AI agents.

But what actually makes something an agent?

Is ChatGPT an agent? Is a deterministic workflow an agent? What happens between an LLM deciding to call a tool and that tool actually executing? Where does memory live? Who controls the loop? Who decides when the agent stops? And what happens when an agent makes the wrong change?

This repository answers those questions from first principles.

No magic. No framework-first abstractions.


What about AI agents confuses you?

Have a question or a concept you want demystified? Check out QUESTIONS.md or submit a question via GitHub Issues. Questions from the community directly shape upcoming modules in this series!


Where Should You Start?


5-Minute Zero-Dependency Quick Start

Clone and run the complete agent loop immediately with Python 3.11+:

git clone https://github.com/tradertanmay/ai-agents-zero-to-hero.git
cd ai-agents-zero-to-hero

# 1. Compare Chatbot vs Workflow vs Agent
python3 01-what-is-an-agent/example.py

# 2. Run the pure Python Observe-Decide-Act loop
python3 02-agent-loop/example.py

# 3. Run your first complete multi-step agent
python3 04-build-your-first-agent/example.py

No API key. No framework. No dependencies. Just Python.


What This Repository Is

A practical, progressive, code-first curriculum designed to take you from:

"I understand LLMs and APIs, but I don't really understand what people mean by an AI agent."

to:

"I understand how agents work internally and can build, debug, evaluate, and reason about production agent systems."


Who This Is For

This repository is built for everyone who wants to understand and build AI agents — whether you are a software engineer, technical lead, researcher, product builder, student, or curious developer.

It is especially designed for you if you:

  • Want to move beyond prompt engineering and understand how autonomous agent systems actually work
  • Know basic Python (or can follow readable standard code)
  • Have used ChatGPT or called LLM APIs, but want to see the underlying machinery behind tools, memory, and loops
  • Hear industry buzzwords like ReAct, Function Calling, Memory, Agent Harness, Multi-Agent, MCP and want a crystal-clear, framework-independent mental model

Core Philosophy: The Model is Not the Agent

A common beginner assumption is:

$$\text{Agent} \stackrel{?}{=} \text{LLM} + \text{Prompt}$$

In reality, a base LLM inference call does not itself maintain persistent application state across turns. An Agent System is a composite computational system comprising a model policy, an execution control loop, state management, and tools, which observes and acts upon an external Environment.

$$\mathbf{Agent\ System = Model/Policy + Runtime/Control\ Loop + State + Tools}$$

flowchart TD
    subgraph AgentSystem["AGENT SYSTEM"]
        M["Model / Policy (Decision Engine)"]
        R["Runtime / Control Loop (Supervisor)"]
        S["State & Memory"]
        T["Tools & Actions"]
        
        R <--> M
        R <--> S
        R <--> T
    end

    AgentSystem <-->|Act / Observe| E["ENVIRONMENT<br/>(APIs, Filesystem, DB, User)"]

    style AgentSystem fill:#f8f9fa,stroke:#333,stroke-width:2px
    style M fill:#f3e5f5,stroke:#7b1fa2
    style R fill:#fff3e0,stroke:#e65100,stroke-width:2px
    style S fill:#ede7f6,stroke:#4527a0
    style T fill:#e8f5e9,stroke:#2e7d32
    style E fill:#e1f5fe,stroke:#0288d1,stroke-width:2px

The Core Agent Loop

At the heart of every agent is the cyclic feedback loop with its environment:

flowchart LR
    O["1. Observe"] --> D["2. Decide"]
    D --> A["3. Act"]
    A --> O

    style O fill:#e1f5fe,stroke:#0288d1
    style D fill:#f3e5f5,stroke:#7b1fa2
    style A fill:#e8f5e9,stroke:#388e3c
  1. Observe: Read the current state of the world, conversation history, and tool feedback.
  2. Decide: The model reasons over the observation and chooses an action or final response.
  3. Act: An execution layer performs the action on the environment.
  4. Observe Again: The environment output becomes a new observation fed back to the model.

Key Rule: The model chooses or requests an action; an execution layer performs it. In a from-scratch agent like the one in this course, that execution layer is our Python runtime. In hosted platforms, the provider may execute certain hosted tools on the application’s behalf.


What People Call an "Agent" — Operational Spectrum

There is no universally agreed boundary across industry and research for what counts as an "agent." In this course, we use the following operational spectrum to make the architectural differences explicit:

flowchart LR
    LLM["1. Raw LLM"] --> Chat["2. Chatbot"]
    Chat --> RAG["3. RAG Pipeline"]
    RAG --> Work["4. Workflow"]
    Work --> ToolLLM["5. Tool-Using LLM"]
    ToolLLM --> Agent["6. Iterative Autonomous Agent"]
    Agent --> Multi["7. Multi-Agent System"]

    style LLM fill:#f5f5f5,stroke:#9e9e9e
    style Chat fill:#e1f5fe,stroke:#0288d1
    style RAG fill:#e0f7fa,stroke:#0097a7
    style Work fill:#fff8e1,stroke:#f57f17
    style ToolLLM fill:#f3e5f5,stroke:#7b1fa2
    style Agent fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px
    style Multi fill:#ede7f6,stroke:#4527a0
StageParadigmHow Decisions Are MadeCan it take actions?Closed Feedback Loop?Agentic characteristics
1Raw LLMPredicts next tokens based on promptNoNoNone
2ChatbotAppends user/assistant turns to historyNoOnly via userLow
3RAG PipelineHardcoded retrieval $\to$ LLM synthesisNoNoLow
4Workflow / ChainFixed code sequence ($A \to B \to C$)Fixed actionsFixed error handlingPredetermined control
5Tool-Using LLMModel selects 1 tool $\to$ returns answerYes 1-shot actionNo iterative retrySome
6Iterative Autonomous AgentDynamic Observe $\to$ Decide $\to$ Act loopYes Dynamic actionsYes Closed runtime feedbackHigh
7Multi-Agent SystemMultiple agents with handoffs & rolesYes Distributed actionsYes Inter-agent feedbackMultiple interacting agents

Conceptual Clarity: Stop Mixing These Up

The 5 Context & Data Concepts

TermWhat It Actually IsLifespan
PromptThe static text template or instruction formatted for a single inference call.Single API call
ContextThe exact collection of tokens passed into the model's attention window at time $t$.Single API call
HistoryThe ordered sequence of prior user/assistant turns and tool observations.Conversation
StateThe full operational data structure (variables, step count, budget, artifacts, scratchpad).Task execution
MemoryKnowledge persisted across tasks/sessions (vector indices, key-value stores, user profiles).Persistent / Multi-session

The 5 System Components

ComponentPrimary ResponsibilityExample
ModelProbabilistic decision-maker and text generator.GPT-4o, Claude 3.5 Sonnet, Gemini 2.0
ToolA callable capability that inspects or mutates the environment.calculator(), read_file(), sql_query()
EnvironmentThe external world the agent observes and acts upon.Filesystem, REST API, Database, Shell
Runtime / HarnessThe supervisor controlling loops, step limits, permissions, and tool execution.Python loop, LangGraph runtime
Agent SystemThe complete composite system (Model + Runtime + Tools + State).Minimal Agent, Coding Assistant

Visual Curriculum Map

Every module is labeled with a difficulty level so you can track your progression:

Level 1: Beginner → Level 2: Builder → Level 3: Systems → Level 4: Production → Level 5: Research
flowchart TD
    subgraph L1["Level 1 — Beginner"]
        M00["00-introduction<br/>(Prerequisites & Setup)"]
        M01["01-what-is-an-agent<br/>(Operational Taxonomy)"]
        M02["02-agent-loop<br/>(Observe-Decide-Act Mechanics)"]
        M00 --> M01 --> M02
    end

    subgraph L2["Level 2 — Builder"]
        M03["03-tools-and-function-calling<br/>(Tool Schemas & MCP)"]
        M04["04-build-your-first-agent<br/>(The Complete End-to-End Agent)"]
        M05["05-state-and-memory<br/>(State, History & Memory Stores)"]
        M02 --> M03 --> M04 --> M05
    end

    subgraph L3["Level 3 — Systems"]
        M06["06-planning-and-reasoning<br/>(ReAct & Task Decomposition)"]
        M07["07-context-engineering<br/>(Window Budgets & Pollution)"]
        M08["08-agent-runtime-and-harness<br/>(Middleware, Budgets & Limits)"]
        M09["09-multi-agent-systems<br/>(Handoffs & Supervisor Patterns)"]
        M05 --> M06 --> M07 --> M08 --> M09
    end

    subgraph L4["Level 4 — Production"]
        M10["10-agent-failures<br/>(Taxonomy & Defense Patterns)"]
        M11["11-agent-evaluation<br/>(Trajectory Evals & Benchmarks)"]
        M12["12-agent-safety-and-verification<br/>(Sandboxing & Approvals)"]
        M13["13-production-agents<br/>(Observability & Persistence)"]
        M09 --> M10 --> M11 --> M12 --> M13
    end

    subgraph L5["Level 5 — Research"]
        M14["14-coding-agents<br/>(Repo Search, Patching & Evals)"]
        M15["15-self-improving-agents<br/>(Meta-Learning & Evolution)"]
        M13 --> M14 --> M15
    end

    style L1 fill:#e8f5e9,stroke:#2e7d32
    style L2 fill:#e1f5fe,stroke:#0288d1
    style L3 fill:#fff8e1,stroke:#f57f17
    style L4 fill:#fbe9e7,stroke:#d84315
    style L5 fill:#f3e5f5,stroke:#6a1b9a

Curriculum Table of Contents (Numerical Progression)

ModuleLevelStatusWhat You Will Learn
00-introductionLevel 1ReadyCourse architecture, prerequisites, mental models
01-what-is-an-agentLevel 1ReadyLLM vs Chatbot vs Workflow vs Agent; operational taxonomy
02-agent-loopLevel 1ReadyThe cyclic Observe-Decide-Act execution loop
03-tools-and-function-callingLevel 2ReadyTool lifecycle, schemas, validation & MCP Deep Dive
04-build-your-first-agentLevel 2ReadyAssembling the first complete agent + 10-case regression scorecard
05-state-and-memoryLevel 2ReadyWorking state, SQLite persistent memory & history compaction
06-planning-and-reasoningLevel 3ReadyReAct, decomposition, reflection, and limits of reasoning
07-context-engineeringLevel 3ReadyContext budgets, compression, and anti-pollution
08-agent-runtime-and-harnessLevel 3ReadyHarness as the OS: step limits, budgets, middleware & aborts
09-multi-agent-systemsLevel 3ReadySupervisor-worker, handoffs, and when NOT to use multi-agent
10-agent-failuresLevel 4ReadyFailure taxonomy, loops, ambiguous writes & reconciliation
11-agent-evaluationLevel 4ReadyAdvanced trajectory evaluation, multi-dimensional metrics & frozen benchmarks
12-agent-safety-and-verificationLevel 4ReadyCapability gating, blast radius, invariants & postconditions
13-production-agentsLevel 4ReadyDurable queues, worker leases, crash recovery, telemetry & health checks
14-coding-agentsLevel 5ReadyRepo search, AST navigation, sandboxed test execution & patch repair loops
15-self-improving-agentsLevel 5ReadyThe 5 adaptation surfaces, isolated candidate evals, multi-dimensional gates & rollback

Zero Dependencies & Framework Independence

We believe you should understand how agents work even if every agent framework disappeared tomorrow.

  • Zero Mandatory Third-Party Packages: Everything runs on standard Python 3.11+.
  • Zero Required API Keys: All core lessons include deterministic simulators and mock LLMs for 100% offline learning and automated testing.
  • Pluggable Real Providers: Want to use a real model? Drop in your API key for OpenAI, Anthropic, Gemini, or local Ollama instances in examples/minimal_agent/.

Once you master the mechanics in this repo, you will understand how modern tools and frameworks organize these responsibilities:

  • LangGraph — maps many of these concepts into a graph-oriented orchestration/runtime model with explicit state, nodes, durable execution, and human-in-the-loop control.
  • AutoGen — provides abstractions for agents, teams, messaging, and event-driven multi-agent orchestration.
  • OpenAI Agents SDK — provides an agent runtime with a built-in agent loop, function tools, handoffs, guardrails, sessions, and tracing.
  • MCP (Model Context Protocol) — an open protocol for connecting AI applications to external capabilities and context providers. MCP servers can expose tools, resources, and prompts through a standardized protocol.

Repository Structure (Strict Numerical Order)

ai-agents-zero-to-hero/
├── README.md # Main course landing page
├── LICENSE # Proprietary License (Tanmay Sah)
├── CONTRIBUTING.md # Contribution & pedagogical standard
├── ROADMAP.md # Curriculum milestone checklist
├── QUESTIONS.md # Community Q&A hub
├── pyproject.toml # Standard Python packaging config
├── .github/ # Issue templates
│
├── 00-introduction/ # Module 0: Prerequisites & mental models
├── 01-what-is-an-agent/ # Module 1: Operational spectrum & agent architecture
├── 02-agent-loop/ # Module 2: The Observe-Decide-Act loop
├── 03-tools-and-function-calling/ # Module 3: Tool schemas, execution lifecycle & MCP
├── 04-build-your-first-agent/ # Module 4: Assembling your first complete agent
├── 05-state-and-memory/ # Module 5: Working State, SQLite Memory & Compaction
├── 06-planning-and-reasoning/ # Module 6: ReAct, Planning & Dynamic Replanning
├── 07-context-engineering/ # Module 7: Token Budgets & Observation Pruning
├── 08-agent-runtime-and-harness/ # Module 8: The Agent Harness / Operating System
├── 09-multi-agent-systems/ # Module 9: Supervisor-Worker & Review Loops
├── 10-agent-failures/             # Module 10: Failure Taxonomy, Ambiguous Writes & Reconciliation
├── 11-agent-evaluation/            # Module 11: Advanced Trajectory Evaluation & Benchmarks
│   ├── README.md                   # Core guide, 5 evaluation levels & commands
│   ├── concepts.md                 # Comprehensive deep-dive & metric formulas
│   ├── eval_cases.json             # Frozen 20-case benchmark test suite
│   ├── example.py                  # Runnable comparative regression benchmark (V1 vs V2)
│   └── exercise.md                 # Production incident reproduction exercise
├── 12-agent-safety-and-verification/ # Module 12: Capability Gating, Blast Radius & Compensation
│   ├── README.md                   # Core guide, 4 layers, 6 controls & commands
│   ├── concepts.md                 # Deep-dive: blast radius formula, 6 invariants & compensation
│   ├── example.py                  # Runnable adversarial safety suite (40+ attack vectors)
│   └── exercise.md                 # Edit comment capability & compensation exercise
├── 13-production-agents/            # Module 13: Durable Queues, Worker Leases & Telemetry
│   ├── README.md                   # Core guide, 6 production concerns & commands
│   ├── concepts.md                 # Deep-dive: durable state machine, reconciliation & health
│   ├── example.py                  # Runnable demo: worker leases, crash recovery & trace waterfalls
│   └── exercise.md                 # Incident root-cause analysis & graceful SIGTERM drain
├── 14-coding-agents/               # Module 14: Repo Search, AST Navigation & Sandboxed Repair
│   ├── README.md                   # Core guide, 9-stage loop, invariants & scorecard
│   ├── concepts.md                 # Deep-dive: AST slicing, execution sandboxes & test immutability
│   ├── example.py                  # Component walkthrough: AST slicing, safety linting & sandboxed repair
│   ├── exercise.md                 # Call-graph extraction, forbidden-pattern security & modulo repair
│   └── demo_repo/                  # Deliberately broken mini-repository (calculator & parser bugs)
├── 15-self-improving-agents/        # Module 15: Self-Improving Agents & Governance
│   ├── README.md                   # Core guide, 5 adaptation surfaces, architecture & commands
│   ├── concepts.md                 # Deep-dive: evaluator isolation, multi-dimensional gates & rollback
│   ├── example.py                  # Walkthrough: holdout boundary, multi-dimensional evals & rollback
│   ├── exercise.md                 # Eval dataset tamper-proofing, AST scanner & canary rollback
│   └── evals/                      # Quarantined eval datasets (development, regression, holdout)
│
├── examples/
│   ├── self_improving_agent/       # Capstone: Versioned Self-Improvement & Rollback Subsystem
│   │   ├── versions.py             # Version registry, promotion statuses & immutable audit trail
│   │   ├── mutation.py             # 5 adaptation surfaces, risk hierarchy & protected verifier guard
│   │   ├── proposer.py             # Failure log analysis & adaptation proposal (holdout isolated)
│   │   ├── candidate.py            # Candidate workspace cloning & isolated patch application
│   │   ├── baseline.py             # Baseline naive agent & candidate AST-localized agent
│   │   ├── evaluator.py            # Frozen regression & holdout benchmark scoring
│   │   ├── promotion.py            # Multi-dimensional gate (zero tolerance on unsafe) & human approval
│   │   ├── rollback.py             # Production rollback manager & ancestor restoration
│   │   └── main.py                 # Flagship end-to-end runnable demonstration
│   ├── coding_assistant/           # Applied Coding Agent: Sandboxed Repair & Verification Subsystem
│   │   ├── repo.py                 # Repository discovery, language detection & tree mapping
│   │   ├── search.py               # Fast code grep & symbol definition lookup
│   │   ├── ast_tools.py            # AST symbol extraction & targeted function slicing
│   │   ├── patch.py                # Unified diffs, targeted replacements & syntax validation
│   │   ├── sandbox.py              # Isolated tempdir cloning, command execution & repo sync
│   │   ├── verifier.py             # Test execution, traceback parsing & safety linter
│   │   ├── agent.py                # Autonomous coding agent loop & verification scorecard
│   │   └── main.py                 # End-to-end runnable demo on demo_repo
│   ├── minimal_agent/              # Modular, runnable showcase agent
│   │   ├── README.md
│   │   ├── llm.py                  # Pluggable LLM interface (Mock, OpenAI, Anthropic, Gemini, Ollama)
│   │   ├── state.py                # State representation & history
│   │   ├── tools.py                # Tool definitions & registry
│   │   ├── runtime.py              # Step controller & budget enforcement
│   │   ├── agent.py                # Pure agent logic
│   │   └── main.py                 # Runnable demo script
│   └── reddit_comment_agent/       # Capstone: Human-in-the-Loop Reddit Comment Agent (Modules 01-09)
│       ├── README.md               # Architecture, module mapping & user guide
│       ├── mock_reddit.py          # Offline simulated Reddit environment
│       ├── reddit.py               # Reddit client interface (Mock + optional PRAW)
│       ├── state.py                # SQLite persistent memory & duplicate prevention
│       ├── tools.py                # Permission-gated tool registry
│       ├── evaluator.py            # 5-criterion quality scorecard
│       ├── approval.py             # Human-in-the-loop review gate & HMAC tokens
│       ├── safety.py               # 4-layer defense, blast limiter, verifiers & compensation
│       ├── runtime.py              # OS harness & rate limits
│       ├── agent.py                # Central coordinator (observe, decide, act)
│       ├── main.py                 # Standalone runnable demo script
│       ├── evals/                  # Benchmark evaluation harness (Module 11)
│       │   ├── cases.json          # Frozen 20-case test suite
│       │   ├── metrics.py          # Multi-dimensional metric calculations
│       │   ├── judges.py           # Deterministic, Heuristic & LLM Judges
│       │   ├── runner.py           # Benchmark execution harness
│       │   └── report.py           # Regression reporting & delta tables
│       └── production/             # Production subsystem (Module 13)
│           ├── config.py           # Configuration, secrets & redaction
│           ├── jobs.py             # State machine & transition validation
│           ├── checkpoints.py      # Durable SQLite queue & worker leases
│           ├── telemetry.py        # Structured JSONL, metrics & tracer
│           ├── health.py           # Liveness, readiness & dependency health
│           ├── recovery.py         # Crash recovery & post-commit reconciliation
│           └── worker.py           # Worker loop & graceful shutdown
└── tests/                          # Unittest verification suite (88 tests)
    ├── test_module_01.py
    ├── test_module_02.py
    ├── test_module_03.py
    ├── test_module_04.py
    ├── test_module_05.py
    ├── test_module_06.py
    ├── test_module_07.py
    ├── test_module_08.py
    ├── test_module_09.py
    ├── test_module_10.py
    ├── test_module_11.py
    ├── test_module_12.py
    ├── test_module_13.py
    ├── test_minimal_agent.py
    └── test_reddit_comment_agent.py

Running the Tests

Run the zero-dependency test suite using standard Python:

python3 -m unittest discover -s tests -v

License

Copyright (c) 2026 Tanmay Sah. All rights reserved.

This work is published under a Proprietary / Source-Available License. No part of the written curriculum, diagrams, code implementations, exercises, or related expressive materials may be reproduced, distributed, modified, or used for commercial, corporate training, or derivative purposes without prior written permission. See LICENSE for full terms.

agent-runtime
ai-agents
autonomous-agents
first-principles
function-calling
llm
mcp
multi-agent
python
zero-to-hero