NVIDIA Voice Agent in C# - Real-time ASR/TTS/LLM with WebSockets (ported from Python)
C#
1
81 commits
updated Feb 17, 2026
A real-time voice agent built with ASP.NET Core 10 that performs Speech-to-Text (ASR), LLM processing, and Text-to-Speech (TTS) using NVIDIA NIM models via ONNX Runtime and TorchSharp.
Ported from: nvidia-transcribe/scenario5
/ws/voice and /ws/logsThe ASR model now supports arbitrarily long audio through intelligent overlapping chunk processing:
Example: Transcribe a 30-minute podcast in one call:
// Just pass any audio length — chunking happens automatically
var transcript = await asrService.TranscribeAsync(thirtyMinuteAudio);
// Result: Full transcription across multiple chunks, no duplicates
Performance: ~100-300 seconds for 30-minute real-world audio (varies by CPU/GPU).
📖 See Audio Chunking Guide for configuration, troubleshooting, and performance tips.
# Clone and build
git clone https://github.com/elbruno/nvidia-voiceagent-cs.git
cd nvidia-voiceagent-cs
dotnet build
# (First-time) install Python helpers for model preparation
pip install -r scripts/onnx/requirements.txt
# (First-time) prepare Parakeet-TDT ASR model
python scripts/onnx/patch_encoder.py --model-dir model-cache/parakeet-tdt-0.6b/onnx
python scripts/onnx/patch_decoder.py --model-dir model-cache/parakeet-tdt-0.6b/onnx
python scripts/onnx/extract_vocab.py --model-dir model-cache/parakeet-tdt-0.6b
# (Windows) run all ASR prep steps at once
pwsh -File scripts/onnx/prepare-models.ps1 -ModelCachePath model-cache
# Run the app
cd NvidiaVoiceAgent
dotnet run
Open http://localhost:5003 in your browser. The app auto-downloads models into ModelHub:ModelCachePath on first run.
First-time setup: See the full step-by-step guide at
docs/guides/model-preparation.md. No GPU? The app works on CPU and falls back to Mock Mode if models are unavailable.
dotnet test
Tests are configuration-driven. Core tests read tests/NvidiaVoiceAgent.Core.Tests/appsettings.Test.json (mirrors the main app) and skip real-model tests when models are not available.
dotnet test --filter "FullyQualifiedName~MockMode"
dotnet test --filter "FullyQualifiedName~RealModel"
dotnet test --filter "FullyQualifiedName~RealModelIntegrationTests"
See tests/NvidiaVoiceAgent.Core.Tests/README.md for the full test configuration guide.
| Model | Type | Size | Status | Notes |
|---|---|---|---|---|
| Parakeet-TDT-0.6B-V2 | ASR | ~2.5 GB | ✅ Auto-download | ONNX format, GPU/CPU |
| PersonaPlex-7B-v1 | LLM | ~16.7 GB | ✅ Available | Full-duplex speech AI, 18 voices, requires HF token |
| TinyLlama-1.1B | LLM | ~2.0 GB | 🔜 Coming soon | Fallback LLM option |
| FastPitch | TTS | ~80 MB | 🔜 Coming soon | Mel-spectrogram generator |
| HiFiGAN | Vocoder | ~55 MB | 🔜 Coming soon | Neural vocoder |
PersonaPlex-7B-v1 is NVIDIA's state-of-the-art full-duplex speech-to-speech conversational AI:
To use PersonaPlex:
Accept the license at https://huggingface.co/nvidia/personaplex-7b-v1
Generate a HuggingFace token with read access
Add token to appsettings.json:
{
"ModelHub": {
"HuggingFaceToken": "hf_your_token_here"
}
}
Download via the UI or API: POST /api/models/PersonaPlex-7B-v1/download
📖 Detailed Setup Guide: See HuggingFace Token Setup Guide for complete instructions, troubleshooting, and security best practices.
Note: PersonaPlex currently runs in mock mode. TorchSharp integration for actual inference is planned.
Edit NvidiaVoiceAgent/appsettings.json:
{
"ModelHub": {
"AutoDownload": true, // Download models on startup
"UseInt8Quantization": true, // Prefer quantized models
"ModelCachePath": "model-cache", // Local cache directory
"HuggingFaceToken": null // Required for gated models (PersonaPlex)
},
"ModelConfig": {
"UseGpu": true, // CUDA acceleration
"Use4BitQuantization": true, // LLM quantization
"PersonaPlexVoice": "voice_0" // Default PersonaPlex voice (0-17)
},
"DebugMode": {
"Enabled": false, // Record conversations for testing
"AudioLogPath": "logs/audio-debug",
"SaveIncomingAudio": true, // Save user voice
"SaveOutgoingAudio": true, // Save TTS responses
"SaveMetadata": true, // Save conversation metadata
"MaxAgeInDays": 7 // Auto-delete old recordings
}
}
nvidia-voiceagent-cs/
├── NvidiaVoiceAgent/ # ASP.NET Core Web App (UI, WebSockets, endpoints)
├── NvidiaVoiceAgent.Core/ # ML/Audio class library (ASR, AudioProcessor, MelSpectrogram)
├── NvidiaVoiceAgent.ModelHub/ # Model download library (HuggingFace integration)
├── tests/ # xUnit test projects
└── docs/ # Detailed documentation
| Document | Description |
|---|---|
| Architecture | Solution structure, project layers, dependency graph |
| Implementation Details | Voice pipeline, audio processing, ONNX inference, model loading |
| PersonaPlex Integration Plan | Detailed plan for PersonaPlex-7B-v1 implementation |
| Debug Mode Guide | Record conversations for testing and analysis |
| E2E Testing with Recorded Audio | Build automated tests from recorded conversations |
| API Reference | HTTP and WebSocket endpoints with message formats |
| Developer Guide | Coding conventions, adding services, testing patterns |
| Troubleshooting | Common issues and solutions |
git checkout -b feature/my-featureSee the Developer Guide for coding standards.
MIT License — see LICENSE for details.
C#
81.9%
HTML
11.3%
Python
6.2%
NVIDIA Voice Agent in C# - Real-time ASR/TTS/LLM with WebSockets (ported from Python)
C#
1
81 commits
updated Feb 17, 2026
A real-time voice agent built with ASP.NET Core 10 that performs Speech-to-Text (ASR), LLM processing, and Text-to-Speech (TTS) using NVIDIA NIM models via ONNX Runtime and TorchSharp.
Ported from: nvidia-transcribe/scenario5
/ws/voice and /ws/logsThe ASR model now supports arbitrarily long audio through intelligent overlapping chunk processing:
Example: Transcribe a 30-minute podcast in one call:
// Just pass any audio length — chunking happens automatically
var transcript = await asrService.TranscribeAsync(thirtyMinuteAudio);
// Result: Full transcription across multiple chunks, no duplicates
Performance: ~100-300 seconds for 30-minute real-world audio (varies by CPU/GPU).
📖 See Audio Chunking Guide for configuration, troubleshooting, and performance tips.
# Clone and build
git clone https://github.com/elbruno/nvidia-voiceagent-cs.git
cd nvidia-voiceagent-cs
dotnet build
# (First-time) install Python helpers for model preparation
pip install -r scripts/onnx/requirements.txt
# (First-time) prepare Parakeet-TDT ASR model
python scripts/onnx/patch_encoder.py --model-dir model-cache/parakeet-tdt-0.6b/onnx
python scripts/onnx/patch_decoder.py --model-dir model-cache/parakeet-tdt-0.6b/onnx
python scripts/onnx/extract_vocab.py --model-dir model-cache/parakeet-tdt-0.6b
# (Windows) run all ASR prep steps at once
pwsh -File scripts/onnx/prepare-models.ps1 -ModelCachePath model-cache
# Run the app
cd NvidiaVoiceAgent
dotnet run
Open http://localhost:5003 in your browser. The app auto-downloads models into ModelHub:ModelCachePath on first run.
First-time setup: See the full step-by-step guide at
docs/guides/model-preparation.md. No GPU? The app works on CPU and falls back to Mock Mode if models are unavailable.
dotnet test
Tests are configuration-driven. Core tests read tests/NvidiaVoiceAgent.Core.Tests/appsettings.Test.json (mirrors the main app) and skip real-model tests when models are not available.
dotnet test --filter "FullyQualifiedName~MockMode"
dotnet test --filter "FullyQualifiedName~RealModel"
dotnet test --filter "FullyQualifiedName~RealModelIntegrationTests"
See tests/NvidiaVoiceAgent.Core.Tests/README.md for the full test configuration guide.
| Model | Type | Size | Status | Notes |
|---|---|---|---|---|
| Parakeet-TDT-0.6B-V2 | ASR | ~2.5 GB | ✅ Auto-download | ONNX format, GPU/CPU |
| PersonaPlex-7B-v1 | LLM | ~16.7 GB | ✅ Available | Full-duplex speech AI, 18 voices, requires HF token |
| TinyLlama-1.1B | LLM | ~2.0 GB | 🔜 Coming soon | Fallback LLM option |
| FastPitch | TTS | ~80 MB | 🔜 Coming soon | Mel-spectrogram generator |
| HiFiGAN | Vocoder | ~55 MB | 🔜 Coming soon | Neural vocoder |
PersonaPlex-7B-v1 is NVIDIA's state-of-the-art full-duplex speech-to-speech conversational AI:
To use PersonaPlex:
Accept the license at https://huggingface.co/nvidia/personaplex-7b-v1
Generate a HuggingFace token with read access
Add token to appsettings.json:
{
"ModelHub": {
"HuggingFaceToken": "hf_your_token_here"
}
}
Download via the UI or API: POST /api/models/PersonaPlex-7B-v1/download
📖 Detailed Setup Guide: See HuggingFace Token Setup Guide for complete instructions, troubleshooting, and security best practices.
Note: PersonaPlex currently runs in mock mode. TorchSharp integration for actual inference is planned.
Edit NvidiaVoiceAgent/appsettings.json:
{
"ModelHub": {
"AutoDownload": true, // Download models on startup
"UseInt8Quantization": true, // Prefer quantized models
"ModelCachePath": "model-cache", // Local cache directory
"HuggingFaceToken": null // Required for gated models (PersonaPlex)
},
"ModelConfig": {
"UseGpu": true, // CUDA acceleration
"Use4BitQuantization": true, // LLM quantization
"PersonaPlexVoice": "voice_0" // Default PersonaPlex voice (0-17)
},
"DebugMode": {
"Enabled": false, // Record conversations for testing
"AudioLogPath": "logs/audio-debug",
"SaveIncomingAudio": true, // Save user voice
"SaveOutgoingAudio": true, // Save TTS responses
"SaveMetadata": true, // Save conversation metadata
"MaxAgeInDays": 7 // Auto-delete old recordings
}
}
nvidia-voiceagent-cs/
├── NvidiaVoiceAgent/ # ASP.NET Core Web App (UI, WebSockets, endpoints)
├── NvidiaVoiceAgent.Core/ # ML/Audio class library (ASR, AudioProcessor, MelSpectrogram)
├── NvidiaVoiceAgent.ModelHub/ # Model download library (HuggingFace integration)
├── tests/ # xUnit test projects
└── docs/ # Detailed documentation
| Document | Description |
|---|---|
| Architecture | Solution structure, project layers, dependency graph |
| Implementation Details | Voice pipeline, audio processing, ONNX inference, model loading |
| PersonaPlex Integration Plan | Detailed plan for PersonaPlex-7B-v1 implementation |
| Debug Mode Guide | Record conversations for testing and analysis |
| E2E Testing with Recorded Audio | Build automated tests from recorded conversations |
| API Reference | HTTP and WebSocket endpoints with message formats |
| Developer Guide | Coding conventions, adding services, testing patterns |
| Troubleshooting | Common issues and solutions |
git checkout -b feature/my-featureSee the Developer Guide for coding standards.
MIT License — see LICENSE for details.
C#
81.9%
HTML
11.3%
Python
6.2%