Companion code for Building Production Voice AI Agents (Book 5, Production AI Agent Engineering series)
Python
0
62 commits
updated Sep 27, 2026
Two voice agents answer the same customer question. One replies in about a second and a half. The other leaves six seconds of silence, long enough that a real caller would wonder if the line dropped. This repository is where that number, and every other number in Building Production Voice AI Agents, actually comes from: a real command, a real recorded run, no invented data.
Companion code for Building Production Voice AI Agents (Book 5 of the Production AI Agent Engineering series). Every number in the book comes from a command in this repository, and the recorded runs ship with it, so report.py reproduces the book's tables with no key and no server.
Each chapter's code is at the tag chNN-end, for example git checkout ch01-end.
📘 Amazon (Kindle) · Paperback in review, link coming soon
Building Production Voice AI Agents — real-time speech, SIP telephony, multi-agent handoffs that survive being spoken instead of typed, an evaluation harness with a leaderboard that refuses to declare a winner without evidence, and a production deployment with a live call bridged from Germany to Kenya to prove it.
| # | Book | Code |
|---|---|---|
| 1 | Agentic Systems Engineering | Agentic-Book |
| 2 | Building Reliable AI Agents | reliable-agents-labs |
| 3 | Production AI Products | triage-app · pkgintel-app · reorder-app |
| 4 | Evaluating AI Agents | agent-evals |
| 5 | Building Production Voice AI Agents | this repo |
uv sync
uv run report.py runs/ch01-*
uv run timeline.py runs/ch01-cascaded-split
uv run pytest -q
You need Docker, uv and a Gemini API key from Google AI Studio.
cp .env.example .env # then paste your key into .env
docker run -d --name livekit -p 7880:7880 -p 7881:7881 -p 7882:7882/udp \
livekit/livekit-server:latest --dev --bind 0.0.0.0 --node-ip 127.0.0.1
uv run realtime_agent.py start # terminal 1, leave running
uv run caller.py --agent realtime --calls 10 --run runs/mine-realtime
uv run report.py runs/mine-realtime
For the cascaded agent, start uv run cascaded_agent.py start instead and call it with --agent cascaded. Settings are environment variables documented at the top of each file.
| File | What it does |
|---|---|
realtime_agent.py | One speech-to-speech model hears and answers (gemini-3.8-live). |
cascaded_agent.py | Speech to text, a text model, then text to speech. |
caller.py | A synthetic caller: plays audio/refund_question.wav in real time and measures time to first audio. |
timeline.py | Lists what the cascaded agent heard, sent to the model and spoke in each call, in order. |
report.py | Summarizes runs: answered calls with a 95% interval, time to first audio with a bootstrap interval for the median, and stage timings. |
web_server.py, web/ | Chapter 3: a page for talking to an agent from a browser, and the small server that gives it a room. |
browser_caller.py | Chapter 3: drives a real Chromium with the recorded question as its microphone (uv run --group browser ...). |
echo_agent.py | Chapter 3: an agent that only sends back what it hears, to measure the transport with no model. |
pauses.py | Chapter 6: writes the question with a longer or shorter pause in the middle, taking the room tone from the pause itself. |
endpoint.py | Chapter 6: counts the turns each question was split into and when the whole question reached the model; --stops and --events show one run's detail. |
talker_agent.py | Chapter 7: an agent that says one recorded answer, so barge-in can be measured with no model and no speech service. |
say.py | Chapter 7: records one spoken line with the speech model and keeps it as a WAV file. |
bargein.py | Chapter 7: how long the agent kept talking after the caller cut in, with --sweep and --anatomy. |
compare.py | Chapter 8: cascaded against realtime on the same tasks, with --answers and --bill. |
tools_agent.py | Chapter 9: the cascaded agent with one tool that books a callback, with CONFIRM=on to read the request back first. |
actions.py | Chapter 9: what the agent actually did per call, with --detail. |
splice.py | Chapter 9: joins recordings with a pause of room tone, to split one request into two turns. |
waiting.py | Chapter 10: what the caller heard while a tool was running, with --gone for callers who left. |
rag_agent.py | Chapter 11: the agent with a look_up tool over the passages in voicelab/knowledge.py. |
answers.py | Chapter 11: what the agent said and what it was given, with --said. |
memory_agent.py | Chapter 12: the agent with a store of named facts, MEMORY and HOLD_OPEN. |
reconnect.py | Chapter 12: a caller whose connection drops mid-call and rejoins the same room. |
kept.py | Chapter 12: what each call's prompt knew, with --reconnect and --store. |
voicelab/ | Settings, run records and statistics shared by the scripts. |
runs/ | Recorded runs used in the book. |
audio/refund_question.wav is a 6.8 second question made with Windows' built-in speech synthesis, 16 kHz mono. audio/agent_answer.wav, audio/interruption.wav and audio/backchannel.wav were each made by one run of say.py and are replayed by Chapter 7, so its experiments cost nothing to repeat.
Calls use Google's paid or free tier depending on your key. The key used for the book allowed 100 requests a day to each text-to-speech model; a cascaded call makes about two. Check your own limits in Google AI Studio.
MIT, see LICENSE.
62 commits
Python
89.0%
TypeScript
7.9%
CSS
1.1%
Companion code for Building Production Voice AI Agents (Book 5, Production AI Agent Engineering series)
Python
0
62 commits
updated Sep 27, 2026
Two voice agents answer the same customer question. One replies in about a second and a half. The other leaves six seconds of silence, long enough that a real caller would wonder if the line dropped. This repository is where that number, and every other number in Building Production Voice AI Agents, actually comes from: a real command, a real recorded run, no invented data.
Companion code for Building Production Voice AI Agents (Book 5 of the Production AI Agent Engineering series). Every number in the book comes from a command in this repository, and the recorded runs ship with it, so report.py reproduces the book's tables with no key and no server.
Each chapter's code is at the tag chNN-end, for example git checkout ch01-end.
📘 Amazon (Kindle) · Paperback in review, link coming soon
Building Production Voice AI Agents — real-time speech, SIP telephony, multi-agent handoffs that survive being spoken instead of typed, an evaluation harness with a leaderboard that refuses to declare a winner without evidence, and a production deployment with a live call bridged from Germany to Kenya to prove it.
| # | Book | Code |
|---|---|---|
| 1 | Agentic Systems Engineering | Agentic-Book |
| 2 | Building Reliable AI Agents | reliable-agents-labs |
| 3 | Production AI Products | triage-app · pkgintel-app · reorder-app |
| 4 | Evaluating AI Agents | agent-evals |
| 5 | Building Production Voice AI Agents | this repo |
uv sync
uv run report.py runs/ch01-*
uv run timeline.py runs/ch01-cascaded-split
uv run pytest -q
You need Docker, uv and a Gemini API key from Google AI Studio.
cp .env.example .env # then paste your key into .env
docker run -d --name livekit -p 7880:7880 -p 7881:7881 -p 7882:7882/udp \
livekit/livekit-server:latest --dev --bind 0.0.0.0 --node-ip 127.0.0.1
uv run realtime_agent.py start # terminal 1, leave running
uv run caller.py --agent realtime --calls 10 --run runs/mine-realtime
uv run report.py runs/mine-realtime
For the cascaded agent, start uv run cascaded_agent.py start instead and call it with --agent cascaded. Settings are environment variables documented at the top of each file.
| File | What it does |
|---|---|
realtime_agent.py | One speech-to-speech model hears and answers (gemini-3.8-live). |
cascaded_agent.py | Speech to text, a text model, then text to speech. |
caller.py | A synthetic caller: plays audio/refund_question.wav in real time and measures time to first audio. |
timeline.py | Lists what the cascaded agent heard, sent to the model and spoke in each call, in order. |
report.py | Summarizes runs: answered calls with a 95% interval, time to first audio with a bootstrap interval for the median, and stage timings. |
web_server.py, web/ | Chapter 3: a page for talking to an agent from a browser, and the small server that gives it a room. |
browser_caller.py | Chapter 3: drives a real Chromium with the recorded question as its microphone (uv run --group browser ...). |
echo_agent.py | Chapter 3: an agent that only sends back what it hears, to measure the transport with no model. |
pauses.py | Chapter 6: writes the question with a longer or shorter pause in the middle, taking the room tone from the pause itself. |
endpoint.py | Chapter 6: counts the turns each question was split into and when the whole question reached the model; --stops and --events show one run's detail. |
talker_agent.py | Chapter 7: an agent that says one recorded answer, so barge-in can be measured with no model and no speech service. |
say.py | Chapter 7: records one spoken line with the speech model and keeps it as a WAV file. |
bargein.py | Chapter 7: how long the agent kept talking after the caller cut in, with --sweep and --anatomy. |
compare.py | Chapter 8: cascaded against realtime on the same tasks, with --answers and --bill. |
tools_agent.py | Chapter 9: the cascaded agent with one tool that books a callback, with CONFIRM=on to read the request back first. |
actions.py | Chapter 9: what the agent actually did per call, with --detail. |
splice.py | Chapter 9: joins recordings with a pause of room tone, to split one request into two turns. |
waiting.py | Chapter 10: what the caller heard while a tool was running, with --gone for callers who left. |
rag_agent.py | Chapter 11: the agent with a look_up tool over the passages in voicelab/knowledge.py. |
answers.py | Chapter 11: what the agent said and what it was given, with --said. |
memory_agent.py | Chapter 12: the agent with a store of named facts, MEMORY and HOLD_OPEN. |
reconnect.py | Chapter 12: a caller whose connection drops mid-call and rejoins the same room. |
kept.py | Chapter 12: what each call's prompt knew, with --reconnect and --store. |
voicelab/ | Settings, run records and statistics shared by the scripts. |
runs/ | Recorded runs used in the book. |
audio/refund_question.wav is a 6.8 second question made with Windows' built-in speech synthesis, 16 kHz mono. audio/agent_answer.wav, audio/interruption.wav and audio/backchannel.wav were each made by one run of say.py and are replayed by Chapter 7, so its experiments cost nothing to repeat.
Calls use Google's paid or free tier depending on your key. The key used for the book allowed 100 requests a day to each text-to-speech model; a cascaded call makes about two. Check your own limits in Google AI Studio.
MIT, see LICENSE.
62 commits
Python
89.0%
TypeScript
7.9%
CSS
1.1%