Have you ever wondered how ChatGPT generates responses? How a model with billions of parameters actually runs on your GPU? Or why inference optimization matters so much?
AIOS is a hands-on learning project where you build an LLM inference engine from scratch. By the end, you won't just use LLMs—you'll understand how inference works under the hood.
By completing this project, you will:
Traditional computing has an OS that bridges applications and hardware. AI computing has an inference engine that bridges LLMs and GPUs:
| Operating System | Inference Engine |
|---|---|
| Process scheduling | Request batching & scheduling |
| Memory management (virtual memory, paging) | KV cache management (paged attention) |
| I/O scheduling | Prefill/decode scheduling |
| Device drivers | Kernel/operator optimization |
| Multi-core parallelism | Multi-GPU tensor parallelism |
Each lesson adds one major optimization. Here's the throughput progression:
| Lesson | Throughput (32 req) | Key Optimization | Multiplier |
|---|---|---|---|
| 3 | ~5 tok/s | Baseline (no KV cache) | 1x |
| 4 | ~25 tok/s | KV cache reuse | 5x |
| 5 | ~30 tok/s | Pre-allocated cache | 1.2x |
| 6 | ~30 tok/s | Paged cache (memory efficiency) | — |
| 7 | ~400 tok/s | Batching | 13x |
| 8 | ~600 tok/s | Continuous batching | 1.5x |
| 9 | ~1000 tok/s | Fused layers | 1.1x |
| 10 | ~1200 tok/s | CUDA graphs | 1.2x |
| 11 | ~1200 tok/s | Sampling (quality) | — |
| 12 | Multi-GPU | Tensor parallelism | model capacity + scale-out |
| 13 | Planned | Prefix caching | prefill savings |
| 14 | Production API | Serving layer | — |
aios/
├── config.py # Engine configuration
├── sampling_params.py # Per-request sampling parameters
├── llm.py # User-facing API
├── engine/
│ ├── llm_engine.py # Orchestrator (model + scheduler + tokenizer)
│ ├── scheduler.py # Prefill-first continuous batching scheduler
│ ├── model_runner.py # GPU execution, KV cache, CUDA graphs, TP
│ ├── sequence.py # Per-request state machine
│ └── block_manager.py # Paged KV cache + prefix caching
├── models/
│ └── qwen3.py # Inference-only Qwen3 (no HF model deps)
├── layers/
│ ├── attention.py # FlashAttention + Triton KV write
│ ├── linear.py # TP-aware linear layers
│ ├── layernorm.py # RMSNorm with fused residual add
│ ├── rotary_embedding.py # Precomputed RoPE
│ ├── activation.py # SiluAndMul (fused gate*up)
│ ├── embed_head.py # Vocab-parallel embedding + LM head
│ └── sampler.py # Gumbel-max sampling
└── utils/
├── loader.py # Weight loader (safetensors → fused modules)
└── context.py # ThreadLocal context for attention metadata
Supports all Qwen3 sizes. Exercises are model-size agnostic — config is loaded from config.json:
| Model | Parameters | Recommended Use |
|---|---|---|
| Qwen3-0.6B | 0.6B | Development, fast iteration |
| Qwen3-1.7B | 1.7B | Development, basic testing |
| Qwen3-4B | 4B | Testing, light benchmarking |
| Qwen3-8B | 8B | Benchmarking |
| Qwen3-14B | 14B | Benchmarking (needs TP=2) |
| Qwen3-32B | 32B | Benchmarking (needs TP=4) |
pip install torch transformers safetensors flash-attn triton xxhash numpy tqdm
# Simple generation
python generate.py --model /path/to/Qwen3-0.6B
# Benchmark throughput
python benchmark.py --model /path/to/Qwen3-0.6B --num-prompts 32
# Using the Python API
python -c "
from aios import LLM, SamplingParams
llm = LLM('/path/to/Qwen3-0.6B')
outputs = llm.generate(['Hello world'], SamplingParams(max_tokens=64))
print(outputs[0]['text'])
"
Start from the beginning:
Each lesson includes:
This project is inspired by:
Educational project. See individual files for details.
46 followers · starred Apr 2026
Python
100.0%
Have you ever wondered how ChatGPT generates responses? How a model with billions of parameters actually runs on your GPU? Or why inference optimization matters so much?
AIOS is a hands-on learning project where you build an LLM inference engine from scratch. By the end, you won't just use LLMs—you'll understand how inference works under the hood.
By completing this project, you will:
Traditional computing has an OS that bridges applications and hardware. AI computing has an inference engine that bridges LLMs and GPUs:
| Operating System | Inference Engine |
|---|---|
| Process scheduling | Request batching & scheduling |
| Memory management (virtual memory, paging) | KV cache management (paged attention) |
| I/O scheduling | Prefill/decode scheduling |
| Device drivers | Kernel/operator optimization |
| Multi-core parallelism | Multi-GPU tensor parallelism |
Each lesson adds one major optimization. Here's the throughput progression:
| Lesson | Throughput (32 req) | Key Optimization | Multiplier |
|---|---|---|---|
| 3 | ~5 tok/s | Baseline (no KV cache) | 1x |
| 4 | ~25 tok/s | KV cache reuse | 5x |
| 5 | ~30 tok/s | Pre-allocated cache | 1.2x |
| 6 | ~30 tok/s | Paged cache (memory efficiency) | — |
| 7 | ~400 tok/s | Batching | 13x |
| 8 | ~600 tok/s | Continuous batching | 1.5x |
| 9 | ~1000 tok/s | Fused layers | 1.1x |
| 10 | ~1200 tok/s | CUDA graphs | 1.2x |
| 11 | ~1200 tok/s | Sampling (quality) | — |
| 12 | Multi-GPU | Tensor parallelism | model capacity + scale-out |
| 13 | Planned | Prefix caching | prefill savings |
| 14 | Production API | Serving layer | — |
aios/
├── config.py # Engine configuration
├── sampling_params.py # Per-request sampling parameters
├── llm.py # User-facing API
├── engine/
│ ├── llm_engine.py # Orchestrator (model + scheduler + tokenizer)
│ ├── scheduler.py # Prefill-first continuous batching scheduler
│ ├── model_runner.py # GPU execution, KV cache, CUDA graphs, TP
│ ├── sequence.py # Per-request state machine
│ └── block_manager.py # Paged KV cache + prefix caching
├── models/
│ └── qwen3.py # Inference-only Qwen3 (no HF model deps)
├── layers/
│ ├── attention.py # FlashAttention + Triton KV write
│ ├── linear.py # TP-aware linear layers
│ ├── layernorm.py # RMSNorm with fused residual add
│ ├── rotary_embedding.py # Precomputed RoPE
│ ├── activation.py # SiluAndMul (fused gate*up)
│ ├── embed_head.py # Vocab-parallel embedding + LM head
│ └── sampler.py # Gumbel-max sampling
└── utils/
├── loader.py # Weight loader (safetensors → fused modules)
└── context.py # ThreadLocal context for attention metadata
Supports all Qwen3 sizes. Exercises are model-size agnostic — config is loaded from config.json:
| Model | Parameters | Recommended Use |
|---|---|---|
| Qwen3-0.6B | 0.6B | Development, fast iteration |
| Qwen3-1.7B | 1.7B | Development, basic testing |
| Qwen3-4B | 4B | Testing, light benchmarking |
| Qwen3-8B | 8B | Benchmarking |
| Qwen3-14B | 14B | Benchmarking (needs TP=2) |
| Qwen3-32B | 32B | Benchmarking (needs TP=4) |
pip install torch transformers safetensors flash-attn triton xxhash numpy tqdm
# Simple generation
python generate.py --model /path/to/Qwen3-0.6B
# Benchmark throughput
python benchmark.py --model /path/to/Qwen3-0.6B --num-prompts 32
# Using the Python API
python -c "
from aios import LLM, SamplingParams
llm = LLM('/path/to/Qwen3-0.6B')
outputs = llm.generate(['Hello world'], SamplingParams(max_tokens=64))
print(outputs[0]['text'])
"
Start from the beginning:
Each lesson includes:
This project is inspired by:
Educational project. See individual files for details.
46 followers · starred Apr 2026
Python
100.0%