A utility for organizing and generating voice packs for Skyrim using TTS
C#
0
218 commits
updated Apr 8, 2026
Generate high-quality custom AI voice acting for Skyrim SE modlists locally. SVGS extracts dialogue from Skyrim plugins, resolves VoiceTypes, and generates speech audio using advanced local AI text-to-speech, and packages the results as installable voicepack mods.
Sound/Voice/) for NPC dialogue, and DBVO voicepacks (Sound/DBVO/) for player dialogue with Dragonborn Voice Over.[angry], [sad], etc.) based on each line's emotion metadata.| Requirement | Details |
|---|---|
| OS | Windows 10 or 11 (64-bit) |
| Runtime | .NET 10 |
| Disk Space | ~4–8 GB for TTS models + space for generated audio |
| Requirement | Details |
|---|---|
| GPU | 6+ GB VRAM (8+ recommended). See backend notes below. |
| Python | 3.10 or newer (virtual environments managed automatically) |
SVGS offers two local Qwen3-TTS backends:
| Requirement | Details |
|---|---|
| Account | ElevenLabs API key (paid, per-character billing) |
| GPU/Python | Not needed |
| Tool | Benefit |
|---|---|
| Creation Kit (free on Steam) | XWMA audio compression (~25 KB vs ~440 KB per line) and LIP mouth animation. Without it: audio still works, files are larger, and NPCs won't move their mouths. |
Python, FFmpeg, and TTS models are all managed automatically by the app.
For a detailed walkthrough, see the Getting Started guide. For a full feature reference, see the User Guide.
Requires .NET 10 SDK.
# Clone the repository
git clone https://github.com/shtaylor/SkyrimVoiceGenStudio.git
cd SkyrimVoiceGenStudio
# Build the solution
dotnet build SkyrimVoiceGenStudio.slnx
# Run the application
dotnet run --project SkyrimVoiceGenStudio/SkyrimVoiceGenStudio.csproj
# Run tests
dotnet test SVGSTests/SVGSTests.csproj
SkyrimVoiceGenStudio/ WPF desktop application (UI, orchestration)
├── Docs/
│ ├── Getting-Started.md First-time setup guide
│ └── User-Guide.md Comprehensive feature reference
├── Views/ WPF pages
├── ViewModels/ MVVM view models
└── Services/ App-level services (server lifecycle, dialogs)
SVGSLib/ Class library (no UI dependency)
├── Models/ EF Core entities and DTOs
├── Providers/ TTS provider abstraction
│ ├── Qwen3/ Local Qwen3-TTS provider
│ └── ElevenLabs/ Cloud ElevenLabs provider
└── Services/ Business logic (import, audio pipeline, export)
SVGSTests/ xUnit test project
PythonServer/ FastAPI TTS server (port 5100)
├── server.py Endpoints: /health, /generate/*, /model/*, /shutdown
└── tts_engine.py Dual-engine TTS wrapper (standard + fast backends)
WhisperServer/ FastAPI Whisper server (port 5101)
├── server.py Endpoints: /health, /transcribe, /shutdown
└── whisper_engine.py faster-whisper transcription wrapper
Dialogue Text → TTS Provider → WAV → LIP Generation → XWMA Encoding → FUZ Packaging
| Stage | With CK Tools | Fallback |
|---|---|---|
| Audio encoding | XWMA via xwmaencode.exe (~25 KB/line) | Raw WAV in FUZ (~440 KB/line) |
| Mouth animation | LIP via LipGenerator.exe | No mouth movement (lipSize=0) |
| FUZ packaging | Built-in (no external tools) | — |
Audio plays correctly in Skyrim in both cases. The Creation Kit is free on Steam.
qwen-tts) + faster-qwen3-tts (CUDA-graph-optimized), PyTorch (CUDA, ROCm, or CPU)Both guides are also accessible from within the app on the Documentation page.
This project is licensed under the GNU General Public License v3.0 (GPL-3.0-or-later).
SVGS uses Mutagen (GPL-3.0) for reading Bethesda plugin files.
A utility for organizing and generating voice packs for Skyrim using TTS
C#
0
218 commits
updated Apr 8, 2026
Generate high-quality custom AI voice acting for Skyrim SE modlists locally. SVGS extracts dialogue from Skyrim plugins, resolves VoiceTypes, and generates speech audio using advanced local AI text-to-speech, and packages the results as installable voicepack mods.
Sound/Voice/) for NPC dialogue, and DBVO voicepacks (Sound/DBVO/) for player dialogue with Dragonborn Voice Over.[angry], [sad], etc.) based on each line's emotion metadata.| Requirement | Details |
|---|---|
| OS | Windows 10 or 11 (64-bit) |
| Runtime | .NET 10 |
| Disk Space | ~4–8 GB for TTS models + space for generated audio |
| Requirement | Details |
|---|---|
| GPU | 6+ GB VRAM (8+ recommended). See backend notes below. |
| Python | 3.10 or newer (virtual environments managed automatically) |
SVGS offers two local Qwen3-TTS backends:
| Requirement | Details |
|---|---|
| Account | ElevenLabs API key (paid, per-character billing) |
| GPU/Python | Not needed |
| Tool | Benefit |
|---|---|
| Creation Kit (free on Steam) | XWMA audio compression (~25 KB vs ~440 KB per line) and LIP mouth animation. Without it: audio still works, files are larger, and NPCs won't move their mouths. |
Python, FFmpeg, and TTS models are all managed automatically by the app.
For a detailed walkthrough, see the Getting Started guide. For a full feature reference, see the User Guide.
Requires .NET 10 SDK.
# Clone the repository
git clone https://github.com/shtaylor/SkyrimVoiceGenStudio.git
cd SkyrimVoiceGenStudio
# Build the solution
dotnet build SkyrimVoiceGenStudio.slnx
# Run the application
dotnet run --project SkyrimVoiceGenStudio/SkyrimVoiceGenStudio.csproj
# Run tests
dotnet test SVGSTests/SVGSTests.csproj
SkyrimVoiceGenStudio/ WPF desktop application (UI, orchestration)
├── Docs/
│ ├── Getting-Started.md First-time setup guide
│ └── User-Guide.md Comprehensive feature reference
├── Views/ WPF pages
├── ViewModels/ MVVM view models
└── Services/ App-level services (server lifecycle, dialogs)
SVGSLib/ Class library (no UI dependency)
├── Models/ EF Core entities and DTOs
├── Providers/ TTS provider abstraction
│ ├── Qwen3/ Local Qwen3-TTS provider
│ └── ElevenLabs/ Cloud ElevenLabs provider
└── Services/ Business logic (import, audio pipeline, export)
SVGSTests/ xUnit test project
PythonServer/ FastAPI TTS server (port 5100)
├── server.py Endpoints: /health, /generate/*, /model/*, /shutdown
└── tts_engine.py Dual-engine TTS wrapper (standard + fast backends)
WhisperServer/ FastAPI Whisper server (port 5101)
├── server.py Endpoints: /health, /transcribe, /shutdown
└── whisper_engine.py faster-whisper transcription wrapper
Dialogue Text → TTS Provider → WAV → LIP Generation → XWMA Encoding → FUZ Packaging
| Stage | With CK Tools | Fallback |
|---|---|---|
| Audio encoding | XWMA via xwmaencode.exe (~25 KB/line) | Raw WAV in FUZ (~440 KB/line) |
| Mouth animation | LIP via LipGenerator.exe | No mouth movement (lipSize=0) |
| FUZ packaging | Built-in (no external tools) | — |
Audio plays correctly in Skyrim in both cases. The Creation Kit is free on Steam.
qwen-tts) + faster-qwen3-tts (CUDA-graph-optimized), PyTorch (CUDA, ROCm, or CPU)Both guides are also accessible from within the app on the Documentation page.
This project is licensed under the GNU General Public License v3.0 (GPL-3.0-or-later).
SVGS uses Mutagen (GPL-3.0) for reading Bethesda plugin files.