A comprehensive collection of atomic Python scripts for learning AI agent development from scratch
Python
21
30 commits
updated Jul 17, 2026
A Python reference for learning and building AI agents from scratch — 208 short scripts covering the stack from API basics to production patterns.
Status: these are my study notes from moving out of academic AI research into AI engineering, published because they might be useful to someone on the same path. They're maintained on a best-effort basis, not actively developed. Corrections and issues are welcome; I may be slow to respond.
Building AI agents requires understanding dozens of interconnected components. Without a map of the landscape, you learn reactively — discovering critical patterns only after committing to an architecture.
This repo is a map of the AI agent landscape — broad rather than deep, by design. Build a mental repertoire upfront so when you design real solutions, you'll recognize which patterns apply and what's in your toolchain. See Scope & Limitations for where the map stops.
Each script isolates one specific concept so you can:
| Phase | Focus | Modules (docs) |
|---|---|---|
| Phase 1: Foundations | OpenAI API + core building blocks | OpenAI API · Pydantic Models · Structured Output · Conversations · Embeddings · Vector Search |
| Phase 2: Core AI Engineering | Daily patterns for building AI systems | Text Preparation · Information Extraction · Classification & Routing · RAG Pipeline · Agent Orchestration · Context Engineering · Memory Patterns · Prompt Engineering · Evaluation |
| Phase 3: Advanced Patterns | Sophisticated agent behaviors + deeper systems | Advanced Agents · Multi-Agent Systems · Advanced Memory · Advanced RAG · Iterative Processing · FastAPI · Clustering & Topics · Evaluation Systems · Document Processing |
| Phase 4: Production & Operations | Taking agents to production | Docker & Containerization · PostgreSQL + pgvector · Observability · Guardrails · Async & Background Jobs · MCP Servers · Cloud Deployment · CI/CD Basics |
| Phase 5: Specialization | Advanced topics for specific use cases | Fine-tuning LLMs · Custom Embeddings · Advanced NLP · Multimodal |
Install dependencies:
git clone https://github.com/karaoglusina/AI-agent-building-blocks.git
cd AI-agent-building-blocks
uv sync --group phase-1 # or --all-groups for everything (several GB: torch, transformers, faiss, bertopic)
source .venv/bin/activate
That's enough for Phase 1. From Phase 2 onwards you also need the spaCy and NLTK data:
uv sync --group phase-2 # phase-2 … phase-5, or --all-groups
python utils/setup_models.py # downloads the spaCy model + NLTK corpora
Most scripts have a TEST_MODE that skips the network call and prints what the script would do, so you can read and run them before deciding to spend anything:
TEST_MODE=1 python "scripts/phase-1-foundations/1.1-openai-basics/01_basic_call.py"
TEST_MODE=1 skips API calls, not imports — you still need the dependencies for that phase installed, but you don't need a key.
# Create .env file
touch .env
# Add your OpenAI API key (replace with your actual key)
echo "OPENAI_API_KEY=sk-..." >> .env
Or manually create .env and add:
OPENAI_API_KEY=sk-your-actual-api-key-here
Then:
python "scripts/phase-1-foundations/1.1-openai-basics/01_basic_call.py"
You are ready to explore, check out modules/ for conceptual walkthroughs and design explanations, structured by phases and modules. You'll find links to related scripts.
python tests/test_all_scripts.py --filter phase-1 # one phase
python tests/test_all_scripts.py # all 208, needs --all-groups installed
This is a smoke test: it runs each script under TEST_MODE=1 and reports what executed, not whether the output is right.
AI-agent-building-blocks/
├── modules/ # Conceptual walkthroughs (docs) - each document links to relevant scripts - Start here!
│ ├── index.md
│ ├── overview.md # The landscape, if you want the map before the code
│ ├── phase-1-foundations/
│ │ ├── 0.0-phase-1-foundations-index.md
│ │ └── 1.1-openai-basics.md
│ │ └── ...
│ └── ...
├── scripts/ # Runnable Python scripts (one concept per file).
│ ├── phase-1-foundations/
│ │ └── 1.1-openai-basics/01_basic_call.py
│ │ └── ...
│ └── ...
├── utils/ # Shared helpers (data loading, env, Phase 4 DB models)
├── tests/
│ └── test_all_scripts.py # Runs every script under TEST_MODE=1
└── data/ # Sample dataset
└── sample_job_data.json
Throughout the docs, a single running agent example is used: AI job market analyzer that;
The dataset (data/sample_job_data.json) contains 1,318 real LinkedIn postings (14 MB) — messy, diverse, and realistic. It's committed directly to the repo, so a plain git clone gets it; no Git LFS or download step needed.
What this is: a map of the components, so you know what exists and roughly what it's for before you commit to an architecture. What it isn't, and where it stops:
If any of these are what you came for, this repo will point you at the right vocabulary and then leave you to it.
Core:
openai - OpenAI API clientpydantic - Data validation and structured outputchromadb - Vector database for embeddingsnumpy - Vector operations and numerical computingNLP & Text Processing:
spacy - NER, POS tagging, lemmatization, dependency parsingnltk - Tokenization, stopwords, sentence segmentationkeybert - Embedding-based keyword extractiontiktoken - Token counting and chunkingrapidfuzz - Fuzzy string matching and similarityProduction & Infrastructure:
fastapi - Modern web framework for APIsuvicorn - ASGI serversqlalchemy - ORM and database toolkitpgvector - PostgreSQL extension for vector similarity searchcelery - Distributed task queue for background jobslangfuse - LLM observability and tracinghttpx - Async HTTP clientML:
bertopic - Topic modeling and clusteringsentence-transformers - Sentence embeddingstransformers - Hugging Face transformers librarypeft - Parameter-efficient fine-tuning (LoRA)torch - PyTorch for deep learningsklearn - Machine learning utilitiesfaiss - Efficient similarity search (Facebook AI)clip - Multimodal embeddings (OpenAI CLIP)PIL (Pillow) - Image processingThe scope and module structure is mainly informed by these great books. I referenced some of the related chapters from the modules.
Among many YouTube videos and blog posts, Dave Ebbelaar's channel significantly influenced the tool and framework choices.
MIT
A comprehensive collection of atomic Python scripts for learning AI agent development from scratch
Python
21
30 commits
updated Jul 17, 2026
A Python reference for learning and building AI agents from scratch — 208 short scripts covering the stack from API basics to production patterns.
Status: these are my study notes from moving out of academic AI research into AI engineering, published because they might be useful to someone on the same path. They're maintained on a best-effort basis, not actively developed. Corrections and issues are welcome; I may be slow to respond.
Building AI agents requires understanding dozens of interconnected components. Without a map of the landscape, you learn reactively — discovering critical patterns only after committing to an architecture.
This repo is a map of the AI agent landscape — broad rather than deep, by design. Build a mental repertoire upfront so when you design real solutions, you'll recognize which patterns apply and what's in your toolchain. See Scope & Limitations for where the map stops.
Each script isolates one specific concept so you can:
| Phase | Focus | Modules (docs) |
|---|---|---|
| Phase 1: Foundations | OpenAI API + core building blocks | OpenAI API · Pydantic Models · Structured Output · Conversations · Embeddings · Vector Search |
| Phase 2: Core AI Engineering | Daily patterns for building AI systems | Text Preparation · Information Extraction · Classification & Routing · RAG Pipeline · Agent Orchestration · Context Engineering · Memory Patterns · Prompt Engineering · Evaluation |
| Phase 3: Advanced Patterns | Sophisticated agent behaviors + deeper systems | Advanced Agents · Multi-Agent Systems · Advanced Memory · Advanced RAG · Iterative Processing · FastAPI · Clustering & Topics · Evaluation Systems · Document Processing |
| Phase 4: Production & Operations | Taking agents to production | Docker & Containerization · PostgreSQL + pgvector · Observability · Guardrails · Async & Background Jobs · MCP Servers · Cloud Deployment · CI/CD Basics |
| Phase 5: Specialization | Advanced topics for specific use cases | Fine-tuning LLMs · Custom Embeddings · Advanced NLP · Multimodal |
Install dependencies:
git clone https://github.com/karaoglusina/AI-agent-building-blocks.git
cd AI-agent-building-blocks
uv sync --group phase-1 # or --all-groups for everything (several GB: torch, transformers, faiss, bertopic)
source .venv/bin/activate
That's enough for Phase 1. From Phase 2 onwards you also need the spaCy and NLTK data:
uv sync --group phase-2 # phase-2 … phase-5, or --all-groups
python utils/setup_models.py # downloads the spaCy model + NLTK corpora
Most scripts have a TEST_MODE that skips the network call and prints what the script would do, so you can read and run them before deciding to spend anything:
TEST_MODE=1 python "scripts/phase-1-foundations/1.1-openai-basics/01_basic_call.py"
TEST_MODE=1 skips API calls, not imports — you still need the dependencies for that phase installed, but you don't need a key.
# Create .env file
touch .env
# Add your OpenAI API key (replace with your actual key)
echo "OPENAI_API_KEY=sk-..." >> .env
Or manually create .env and add:
OPENAI_API_KEY=sk-your-actual-api-key-here
Then:
python "scripts/phase-1-foundations/1.1-openai-basics/01_basic_call.py"
You are ready to explore, check out modules/ for conceptual walkthroughs and design explanations, structured by phases and modules. You'll find links to related scripts.
python tests/test_all_scripts.py --filter phase-1 # one phase
python tests/test_all_scripts.py # all 208, needs --all-groups installed
This is a smoke test: it runs each script under TEST_MODE=1 and reports what executed, not whether the output is right.
AI-agent-building-blocks/
├── modules/ # Conceptual walkthroughs (docs) - each document links to relevant scripts - Start here!
│ ├── index.md
│ ├── overview.md # The landscape, if you want the map before the code
│ ├── phase-1-foundations/
│ │ ├── 0.0-phase-1-foundations-index.md
│ │ └── 1.1-openai-basics.md
│ │ └── ...
│ └── ...
├── scripts/ # Runnable Python scripts (one concept per file).
│ ├── phase-1-foundations/
│ │ └── 1.1-openai-basics/01_basic_call.py
│ │ └── ...
│ └── ...
├── utils/ # Shared helpers (data loading, env, Phase 4 DB models)
├── tests/
│ └── test_all_scripts.py # Runs every script under TEST_MODE=1
└── data/ # Sample dataset
└── sample_job_data.json
Throughout the docs, a single running agent example is used: AI job market analyzer that;
The dataset (data/sample_job_data.json) contains 1,318 real LinkedIn postings (14 MB) — messy, diverse, and realistic. It's committed directly to the repo, so a plain git clone gets it; no Git LFS or download step needed.
What this is: a map of the components, so you know what exists and roughly what it's for before you commit to an architecture. What it isn't, and where it stops:
If any of these are what you came for, this repo will point you at the right vocabulary and then leave you to it.
Core:
openai - OpenAI API clientpydantic - Data validation and structured outputchromadb - Vector database for embeddingsnumpy - Vector operations and numerical computingNLP & Text Processing:
spacy - NER, POS tagging, lemmatization, dependency parsingnltk - Tokenization, stopwords, sentence segmentationkeybert - Embedding-based keyword extractiontiktoken - Token counting and chunkingrapidfuzz - Fuzzy string matching and similarityProduction & Infrastructure:
fastapi - Modern web framework for APIsuvicorn - ASGI serversqlalchemy - ORM and database toolkitpgvector - PostgreSQL extension for vector similarity searchcelery - Distributed task queue for background jobslangfuse - LLM observability and tracinghttpx - Async HTTP clientML:
bertopic - Topic modeling and clusteringsentence-transformers - Sentence embeddingstransformers - Hugging Face transformers librarypeft - Parameter-efficient fine-tuning (LoRA)torch - PyTorch for deep learningsklearn - Machine learning utilitiesfaiss - Efficient similarity search (Facebook AI)clip - Multimodal embeddings (OpenAI CLIP)PIL (Pillow) - Image processingThe scope and module structure is mainly informed by these great books. I referenced some of the related chapters from the modules.
Among many YouTube videos and blog posts, Dave Ebbelaar's channel significantly influenced the tool and framework choices.
MIT