Windows desktop app for voice cloning with Chatterbox TTS: record or import a voice, write a (multi-speaker) script, render a WAV.
Python
1
17 commits
updated Sep 30, 2026
A small Windows desktop app for voice cloning with Chatterbox TTS by Resemble AI.
[laugh]-style sound tags..wav file.
Everything runs locally on your PC. Step-by-step guide with screenshots: HOWTO.md.
The name comes from the lyrebird, an Australian bird that can copy almost any sound it hears.
Lyrebird-win64.zip (about 33 MB) from the Releases page and unzip it anywhere.Lyrebird\Lyrebird.exe. The exe isn't code-signed yet, so if Windows SmartScreen appears, click More info → Run anyway.%LOCALAPPDATA%\Lyrebird. That's about a 3 GB download with an NVIDIA GPU, and much less without one. It needs about 6 GB of free disk space once unpacked. The window shows the current step, how much is on disk so far and how long it has taken. Setup took under 2 minutes on a fast connection. Later launches start straight away.Requirements: 64-bit Windows 10 or 11 and an internet connection for the first run. An NVIDIA GPU is strongly recommended: about 6 GB of VRAM, since the models use 3 to 4 GB while rendering. Setup picks the PyTorch build that matches your NVIDIA driver. Without an NVIDIA GPU, Lyrebird installs the CPU build and runs on the CPU. That works, but slowly: expect tens of seconds or more per sentence.
Documents\Lyrebird\voices. The ··· menu next to the voice renames and deletes voices.Alice: Did you hear the news?
Bob: No, what happened?
Built-in: I'm the narrator.
A line only switches speaker when the name matches a saved voice (ignoring capitals). Built-in is the model's own voice. Lines without a name continue with the current speaker, and text before the first name uses the voice selected in the Voice list.[laugh], [sigh] and [whispering]. They're highlighted blue when the selected model supports them. With other models they turn amber and are left out when rendering, so they aren't read aloud. Unknown tags are underlined in red..lyrebird file, and Open... (Ctrl+O) brings it all back. It also opens .txt and .md scripts.| Option | Range (default) | Models | Effect |
|---|---|---|---|
| Voice | saved voices or (Built-in voice) | all | Who speaks. It's also the speaker for any text before the first Name: line. The ··· menu renames or deletes it, or opens Voice settings to give it its own slider values. |
| Source | microphones and "What you hear" outputs (Windows default microphone) | all | Used by Record 10 s. Microphones record at the device's own sample rate and keep the loudest channel. If a mic won't open through WASAPI, Lyrebird retries it through MME and DirectSound. "What you hear" records the output device at 48 kHz and mixes it to mono. A silent recording is rejected rather than saved as a voice. |
| Model | Standard / Turbo / Multilingual (Standard) | all | Standard: the original 0.5B English model. Best quality, and supports every option. Turbo: a smaller 350M English model that's much faster and understands sound tags. Multilingual: the 0.5B model for 23 languages. Only one model is kept in memory; switching models unloads the previous one. |
| Language | 23 languages (en) | Multilingual | The language of the text. The voice sample can be in any language, but one in the same language as the text sounds most natural. Supported: Arabic, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Italian, Japanese, Korean, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Turkish. |
| Exaggeration | 0.25 – 2.0 (0.5) | Standard, Multilingual | Emotional intensity. At 0.5 the delivery is neutral. At 0.7 and above it becomes more dramatic and tends to speed up. Very high values can become unstable. |
| CFG / pace | 0.0 – 1.0 (0.5) | Standard, Multilingual | Classifier-free guidance weight, i.e. how closely the output follows the voice sample's style. Lower values give slower, more deliberate speech. Try about 0.3 when you raise Exaggeration, or when the voice sample talks fast. With Multilingual, set it to 0 when the voice sample is in a different language from the text, to reduce accent carry-over. On Standard, 0 is treated as 0.001 because Chatterbox 0.1.7 crashes at exactly 0. |
| Temperature | 0.05 – 2.0 (0.8) | all | Sampling randomness. Lower values give flatter, more consistent speech. Higher values give more variety but also more mistakes. |
| Seed | 0 – 2³¹−1 (0) | all | 0 gives a different take on every render. Any other number makes a render repeatable: same seed, text, voice, settings and hardware give the same result. |
| Save to | folder (Documents\Lyrebird) | all | Where renders are written, as lyrebird_<YYYYmmdd_HHMMSS>.wav (or .flac). The folder is created if it doesn't exist. |
| Format | WAV 16-bit / WAV 32-bit float / FLAC 24-bit (WAV 16-bit) | all | File type of the render. Play output plays 16-bit WAV directly; other formats open in your default player. |
| Sample rate | 24 kHz (model native) / 44.1 kHz / 48 kHz (24 kHz) | all | The models produce 24 kHz. The other rates resample the result, which doesn't add detail but saves converting it later. |
| Voice settings | per voice: Exaggeration, CFG / pace, Temperature (off) | all | Settings switched on replace the main sliders whenever that voice speaks. They're saved as voices\<name>.json next to the sample. Turbo only uses Temperature. |
Exaggeration and CFG are greyed out for Turbo because the Turbo model doesn't support them. Language is greyed out except for Multilingual.
Output is mono. Every setting above is saved when you render or close the window, and restored on the next launch.
Only the Turbo model supports these tags. They're from the Turbo model's own tokenizer:
[laugh] [chuckle] [sigh] [gasp] [cough] [clear throat] [sniff] [groan] [shush] [whispering] [angry] [happy] [sarcastic] [surprised] [fear] [crying] [dramatic] [narration] [advertisement]
| Path | What's there |
|---|---|
%LOCALAPPDATA%\Lyrebird\ | The Python runtime, PyTorch and Chatterbox (env, python, uv-cache), plus settings.json, setup.log and lyrebird.log. |
Documents\Lyrebird\ | Your renders. Voices are in voices\: one audio file per voice, plus <name>.json if the voice has its own settings. |
%USERPROFILE%\.cache\huggingface\ | Model weights, about 3 to 4 GB per model used. |
%USERPROFILE%\.pkuseg\ | The Chinese word-segmentation model, downloaded from GitHub the first time Multilingual loads. |
Uninstall: delete the unzipped Lyrebird folder and %LOCALAPPDATA%\Lyrebird. To also free the model space, delete the models--ResembleAI--chatterbox* folders in the Hugging Face cache. If you used Multilingual, also delete %USERPROFILE%\.pkuseg. Your voices and renders in Documents\Lyrebird stay unless you delete them.
Network use:
HF_HUB_OFFLINE=1 to stop this.Nothing you record or render leaves your PC.
| Variable | Effect |
|---|---|
HF_HOME | Moves the Hugging Face cache somewhere other than %USERPROFILE%\.cache\huggingface. |
HF_TOKEN | Hugging Face access token. It isn't needed for the public Chatterbox models, but it avoids anonymous rate limits. |
HF_HUB_OFFLINE=1 | Never contacts Hugging Face. The models must already be cached. |
PKUSEG_HOME | Where Multilingual stores its Chinese segmentation model (default %USERPROFILE%\.pkuseg). To use Multilingual offline, this folder must already be there. |
Lyrebird has no console window, so its output goes to log files in %LOCALAPPDATA%\Lyrebird:
setup.log: the first-run setup.lyrebird.log: the app itself. It's overwritten on each launch.Please attach these to bug reports.
Lyrebird.exe is a small launcher (launcher.py) frozen with PyInstaller. It bundles uv along with lyrebird.py and requirements.txt. On first run it does the following:
uv venv --managed-python --python 3.12 downloads a standalone Python that includes tkinter.uv pip install -r requirements.txt --torch-backend auto installs PyTorch 2.7.1 in the CUDA build that matches the NVIDIA driver, or the CPU build if there's no GPU. If no CUDA build fits, for example because the driver is too old, it retries with the CPU build. A network error doesn't trigger this fallback: setup stops and resumes the next time Lyrebird starts, so a GPU PC is never quietly left on the CPU build.uv pip install --no-deps chatterbox-tts==0.1.7. Chatterbox pins torch==2.6.0, which has no RTX 50-series (Blackwell) support, so requirements.txt lists its dependencies directly.A marker file stores a hash of requirements.txt. When a new release changes the dependencies, setup runs again; otherwise the launcher starts the app immediately. This keeps the download at about 33 MB instead of about 5 GB. Almost all of that 5 GB is PyTorch's CUDA libraries, which are now downloaded to the user's PC instead.
You need Windows, Git and uv (winget install astral-sh.uv).
git clone https://github.com/zebadrabbit/Lyrebird.git
cd Lyrebird
powershell -ExecutionPolicy Bypass -File build.ps1
build.ps1 does the following:
.venv with PyInstaller, uv and the packages the tests import (numpy, CustomTkinter, tkinterdnd2).%TEMP%\lyrebird-build. It builds there because OneDrive/Dropbox-synced folders lock freshly written files.dist\Lyrebird-win64.zip.Pushing a v* tag makes GitHub Actions build the zip and attach it to a GitHub Release.
Run from source (uses the same first-run setup as the release):
uv run --no-project --managed-python --python 3.12 --with customtkinter==6.0.0 launcher.py
Tests (no model or GPU needed):
uv run --no-project --managed-python --python 3.12 --with "numpy<2" --with customtkinter==6.0.0 --with tkinterdnd2==0.6.3 test_lyrebird.py
lyrebird.py the app: tkinter GUI + Chatterbox wrapper
launcher.py Lyrebird.exe: first-run setup window, then starts lyrebird.py
test_lyrebird.py tests for the text chunker and dialogue parser
requirements.txt runtime dependencies installed by the launcher
build.ps1 builds dist\Lyrebird-win64.zip
HOWTO.md usage guide with screenshots (docs/images/)
examples/ dialogue scripts to try (Turbo + sound tags)
brand/ logo, wordmarks and app icon (lyrebird.ico, SVG, PNG)
Free code signing provided by SignPath.io, certificate by SignPath Foundation.
Lyrebird.exe in each GitHub release, built by this repository's GitHub Actions workflow from the tagged commit. Bundled upstream files (uv, the Python runtime) are included as their authors published them and aren't re-signed.Lyrebird has no accounts, telemetry or analytics. Your recordings, voices, scripts and renders stay on your PC. It connects to other systems only to download what it needs to run:
HF_HUB_OFFLINE=1 to turn that off.Those services see the requests like any download (your IP address, for example), under their own privacy policies. Nothing you record or type is sent anywhere.
Only clone voices you have permission to use. Every file Chatterbox generates carries Resemble AI's Perth watermark: an imperceptible mark that detection tools can find. Lyrebird doesn't remove it. Don't use this software to impersonate people, commit fraud or deceive anyone.
brand/, see its README): logo and icon are part of Lyrebird (MIT). The fonts in brand/fonts/ (Bricolage Grotesque, Atkinson Hyperlegible Next, IBM Plex Mono) are under the SIL Open Font License 1.1; the license texts are next to them.Python
83.4%
HTML
7.9%
CSS
4.8%
PowerShell
3.9%
Windows desktop app for voice cloning with Chatterbox TTS: record or import a voice, write a (multi-speaker) script, render a WAV.
Python
1
17 commits
updated Sep 30, 2026
A small Windows desktop app for voice cloning with Chatterbox TTS by Resemble AI.
[laugh]-style sound tags..wav file.
Everything runs locally on your PC. Step-by-step guide with screenshots: HOWTO.md.
The name comes from the lyrebird, an Australian bird that can copy almost any sound it hears.
Lyrebird-win64.zip (about 33 MB) from the Releases page and unzip it anywhere.Lyrebird\Lyrebird.exe. The exe isn't code-signed yet, so if Windows SmartScreen appears, click More info → Run anyway.%LOCALAPPDATA%\Lyrebird. That's about a 3 GB download with an NVIDIA GPU, and much less without one. It needs about 6 GB of free disk space once unpacked. The window shows the current step, how much is on disk so far and how long it has taken. Setup took under 2 minutes on a fast connection. Later launches start straight away.Requirements: 64-bit Windows 10 or 11 and an internet connection for the first run. An NVIDIA GPU is strongly recommended: about 6 GB of VRAM, since the models use 3 to 4 GB while rendering. Setup picks the PyTorch build that matches your NVIDIA driver. Without an NVIDIA GPU, Lyrebird installs the CPU build and runs on the CPU. That works, but slowly: expect tens of seconds or more per sentence.
Documents\Lyrebird\voices. The ··· menu next to the voice renames and deletes voices.Alice: Did you hear the news?
Bob: No, what happened?
Built-in: I'm the narrator.
A line only switches speaker when the name matches a saved voice (ignoring capitals). Built-in is the model's own voice. Lines without a name continue with the current speaker, and text before the first name uses the voice selected in the Voice list.[laugh], [sigh] and [whispering]. They're highlighted blue when the selected model supports them. With other models they turn amber and are left out when rendering, so they aren't read aloud. Unknown tags are underlined in red..lyrebird file, and Open... (Ctrl+O) brings it all back. It also opens .txt and .md scripts.| Option | Range (default) | Models | Effect |
|---|---|---|---|
| Voice | saved voices or (Built-in voice) | all | Who speaks. It's also the speaker for any text before the first Name: line. The ··· menu renames or deletes it, or opens Voice settings to give it its own slider values. |
| Source | microphones and "What you hear" outputs (Windows default microphone) | all | Used by Record 10 s. Microphones record at the device's own sample rate and keep the loudest channel. If a mic won't open through WASAPI, Lyrebird retries it through MME and DirectSound. "What you hear" records the output device at 48 kHz and mixes it to mono. A silent recording is rejected rather than saved as a voice. |
| Model | Standard / Turbo / Multilingual (Standard) | all | Standard: the original 0.5B English model. Best quality, and supports every option. Turbo: a smaller 350M English model that's much faster and understands sound tags. Multilingual: the 0.5B model for 23 languages. Only one model is kept in memory; switching models unloads the previous one. |
| Language | 23 languages (en) | Multilingual | The language of the text. The voice sample can be in any language, but one in the same language as the text sounds most natural. Supported: Arabic, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Italian, Japanese, Korean, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Turkish. |
| Exaggeration | 0.25 – 2.0 (0.5) | Standard, Multilingual | Emotional intensity. At 0.5 the delivery is neutral. At 0.7 and above it becomes more dramatic and tends to speed up. Very high values can become unstable. |
| CFG / pace | 0.0 – 1.0 (0.5) | Standard, Multilingual | Classifier-free guidance weight, i.e. how closely the output follows the voice sample's style. Lower values give slower, more deliberate speech. Try about 0.3 when you raise Exaggeration, or when the voice sample talks fast. With Multilingual, set it to 0 when the voice sample is in a different language from the text, to reduce accent carry-over. On Standard, 0 is treated as 0.001 because Chatterbox 0.1.7 crashes at exactly 0. |
| Temperature | 0.05 – 2.0 (0.8) | all | Sampling randomness. Lower values give flatter, more consistent speech. Higher values give more variety but also more mistakes. |
| Seed | 0 – 2³¹−1 (0) | all | 0 gives a different take on every render. Any other number makes a render repeatable: same seed, text, voice, settings and hardware give the same result. |
| Save to | folder (Documents\Lyrebird) | all | Where renders are written, as lyrebird_<YYYYmmdd_HHMMSS>.wav (or .flac). The folder is created if it doesn't exist. |
| Format | WAV 16-bit / WAV 32-bit float / FLAC 24-bit (WAV 16-bit) | all | File type of the render. Play output plays 16-bit WAV directly; other formats open in your default player. |
| Sample rate | 24 kHz (model native) / 44.1 kHz / 48 kHz (24 kHz) | all | The models produce 24 kHz. The other rates resample the result, which doesn't add detail but saves converting it later. |
| Voice settings | per voice: Exaggeration, CFG / pace, Temperature (off) | all | Settings switched on replace the main sliders whenever that voice speaks. They're saved as voices\<name>.json next to the sample. Turbo only uses Temperature. |
Exaggeration and CFG are greyed out for Turbo because the Turbo model doesn't support them. Language is greyed out except for Multilingual.
Output is mono. Every setting above is saved when you render or close the window, and restored on the next launch.
Only the Turbo model supports these tags. They're from the Turbo model's own tokenizer:
[laugh] [chuckle] [sigh] [gasp] [cough] [clear throat] [sniff] [groan] [shush] [whispering] [angry] [happy] [sarcastic] [surprised] [fear] [crying] [dramatic] [narration] [advertisement]
| Path | What's there |
|---|---|
%LOCALAPPDATA%\Lyrebird\ | The Python runtime, PyTorch and Chatterbox (env, python, uv-cache), plus settings.json, setup.log and lyrebird.log. |
Documents\Lyrebird\ | Your renders. Voices are in voices\: one audio file per voice, plus <name>.json if the voice has its own settings. |
%USERPROFILE%\.cache\huggingface\ | Model weights, about 3 to 4 GB per model used. |
%USERPROFILE%\.pkuseg\ | The Chinese word-segmentation model, downloaded from GitHub the first time Multilingual loads. |
Uninstall: delete the unzipped Lyrebird folder and %LOCALAPPDATA%\Lyrebird. To also free the model space, delete the models--ResembleAI--chatterbox* folders in the Hugging Face cache. If you used Multilingual, also delete %USERPROFILE%\.pkuseg. Your voices and renders in Documents\Lyrebird stay unless you delete them.
Network use:
HF_HUB_OFFLINE=1 to stop this.Nothing you record or render leaves your PC.
| Variable | Effect |
|---|---|
HF_HOME | Moves the Hugging Face cache somewhere other than %USERPROFILE%\.cache\huggingface. |
HF_TOKEN | Hugging Face access token. It isn't needed for the public Chatterbox models, but it avoids anonymous rate limits. |
HF_HUB_OFFLINE=1 | Never contacts Hugging Face. The models must already be cached. |
PKUSEG_HOME | Where Multilingual stores its Chinese segmentation model (default %USERPROFILE%\.pkuseg). To use Multilingual offline, this folder must already be there. |
Lyrebird has no console window, so its output goes to log files in %LOCALAPPDATA%\Lyrebird:
setup.log: the first-run setup.lyrebird.log: the app itself. It's overwritten on each launch.Please attach these to bug reports.
Lyrebird.exe is a small launcher (launcher.py) frozen with PyInstaller. It bundles uv along with lyrebird.py and requirements.txt. On first run it does the following:
uv venv --managed-python --python 3.12 downloads a standalone Python that includes tkinter.uv pip install -r requirements.txt --torch-backend auto installs PyTorch 2.7.1 in the CUDA build that matches the NVIDIA driver, or the CPU build if there's no GPU. If no CUDA build fits, for example because the driver is too old, it retries with the CPU build. A network error doesn't trigger this fallback: setup stops and resumes the next time Lyrebird starts, so a GPU PC is never quietly left on the CPU build.uv pip install --no-deps chatterbox-tts==0.1.7. Chatterbox pins torch==2.6.0, which has no RTX 50-series (Blackwell) support, so requirements.txt lists its dependencies directly.A marker file stores a hash of requirements.txt. When a new release changes the dependencies, setup runs again; otherwise the launcher starts the app immediately. This keeps the download at about 33 MB instead of about 5 GB. Almost all of that 5 GB is PyTorch's CUDA libraries, which are now downloaded to the user's PC instead.
You need Windows, Git and uv (winget install astral-sh.uv).
git clone https://github.com/zebadrabbit/Lyrebird.git
cd Lyrebird
powershell -ExecutionPolicy Bypass -File build.ps1
build.ps1 does the following:
.venv with PyInstaller, uv and the packages the tests import (numpy, CustomTkinter, tkinterdnd2).%TEMP%\lyrebird-build. It builds there because OneDrive/Dropbox-synced folders lock freshly written files.dist\Lyrebird-win64.zip.Pushing a v* tag makes GitHub Actions build the zip and attach it to a GitHub Release.
Run from source (uses the same first-run setup as the release):
uv run --no-project --managed-python --python 3.12 --with customtkinter==6.0.0 launcher.py
Tests (no model or GPU needed):
uv run --no-project --managed-python --python 3.12 --with "numpy<2" --with customtkinter==6.0.0 --with tkinterdnd2==0.6.3 test_lyrebird.py
lyrebird.py the app: tkinter GUI + Chatterbox wrapper
launcher.py Lyrebird.exe: first-run setup window, then starts lyrebird.py
test_lyrebird.py tests for the text chunker and dialogue parser
requirements.txt runtime dependencies installed by the launcher
build.ps1 builds dist\Lyrebird-win64.zip
HOWTO.md usage guide with screenshots (docs/images/)
examples/ dialogue scripts to try (Turbo + sound tags)
brand/ logo, wordmarks and app icon (lyrebird.ico, SVG, PNG)
Free code signing provided by SignPath.io, certificate by SignPath Foundation.
Lyrebird.exe in each GitHub release, built by this repository's GitHub Actions workflow from the tagged commit. Bundled upstream files (uv, the Python runtime) are included as their authors published them and aren't re-signed.Lyrebird has no accounts, telemetry or analytics. Your recordings, voices, scripts and renders stay on your PC. It connects to other systems only to download what it needs to run:
HF_HUB_OFFLINE=1 to turn that off.Those services see the requests like any download (your IP address, for example), under their own privacy policies. Nothing you record or type is sent anywhere.
Only clone voices you have permission to use. Every file Chatterbox generates carries Resemble AI's Perth watermark: an imperceptible mark that detection tools can find. Lyrebird doesn't remove it. Don't use this software to impersonate people, commit fraud or deceive anyone.
brand/, see its README): logo and icon are part of Lyrebird (MIT). The fonts in brand/fonts/ (Bricolage Grotesque, Atkinson Hyperlegible Next, IBM Plex Mono) are under the SIL Open Font License 1.1; the license texts are next to them.Python
83.4%
HTML
7.9%
CSS
4.8%
PowerShell
3.9%