SharpAudio is a high-performance Text-to-Speech and Speech-to-Text REST API with OpenAI-compatible endpoints, highly vibe-coded:blush:. Built with .NET 10, it features a Vue.js frontend, Docker containerization, and support for multiple state-of-the-art TTS models (Kokoro, Supertonic-3) and ASR models (Whisper, Nemotron). Which task types a given deployment serves — TTS, ASR, or both — is controlled by a single SERVER_MODE environment variable.
/v1/audio/speech and /v1/audio/transcriptions endpointsSERVER_MODEdocker build --target# Clone the repository
git clone https://github.com/Fhrozen/SharpAudio.git
cd SharpAudio
# Start the service
docker compose up -d
# Access the web interface
open http://localhost:5768
# Or use the API directly
curl -X POST http://localhost:5768/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "kokoro-q4",
"input": "Hello world, this is a test of SharpAudio.",
"voice": "af_bella",
"speed": 1.0
}' \
--output speech.wav
The service will automatically download required models from Hugging Face on first startup.
By default the container only serves TTS. To enable Speech-to-Text as well, set SERVER_MODE:
# tts (default) | asr | both
SERVER_MODE=both docker compose up -d
# Transcribe an audio file (Whisper by default)
curl -X POST http://localhost:5768/v1/audio/transcriptions \
-F file=@sample.wav \
-F model=whisper-base
| Endpoint | Method | Description |
|---|---|---|
/health | GET | Health check endpoint |
/api/server-info | GET | Which task types (TTS/ASR) this server instance was started with |
/api/models | GET | List available TTS models with details (gated by SERVER_MODE) |
/api/asr-models | GET | List available ASR models with details (gated by SERVER_MODE) |
/v1/models | GET | OpenAI-compatible model listing (merges TTS + ASR catalogs) |
/v1/audio/speech | POST | OpenAI-compatible speech synthesis (gated by SERVER_MODE) |
/v1/audio/transcriptions | POST | OpenAI-compatible audio transcription (gated by SERVER_MODE) |
/swagger | GET | Interactive API documentation |
Text-to-Speech
| Model | Description | Languages | Voices | Quality |
|---|---|---|---|---|
kokoro-q4 | Quantized Kokoro 82M | 3+ | 5+ | Fast, low memory |
kokoro-full | Full precision Kokoro 82M | 3+ | 5+ | High quality |
supertonic-3 | Supertonic multilingual | 3+ | 10 styles | Production grade |
Speech-to-Text
| Model | Description | Languages | Notes |
|---|---|---|---|
whisper-base | OpenAI Whisper base (GGML, via whisper.cpp) | Auto-detect + multilingual | Whole-file batch transcription |
nemotron-3.5 | NVIDIA Nemotron 3.5 streaming ASR (FastConformer-RNNT, INT4 ONNX) | 35+ | Cache-aware chunked decoding |
See Models Documentation for detailed information about each model.
English (US/GB), Spanish, French, Hindi, Italian, Japanese, Portuguese (BR), Chinese (CN), Korean, German, Dutch, Arabic, Russian, Turkish, Polish, Swedish, Danish, Norwegian, Finnish, Greek, Czech, Romanian, Hungarian, Thai, Vietnamese, Indonesian, Hebrew, Ukrainian, and more.
Key environment variables:
HTTP_PORT=5768 # HTTP port
SERVER_MODE=tts # tts (default) | asr | both
MODEL_CACHE_DIR=/cache # Model cache directory
ESPEAK_DATA_DIR=/app/assets/espeak-ng-data # espeak-ng data
MODEL_IDLE_TIMEOUT_SECONDS=60 # Model unload timeout
See Configuration Guide for all options.
# Backend
cd src/SharpAudio.Api
dotnet restore
dotnet run
# Frontend
cd frontend
corepack enable pnpm
pnpm install
pnpm dev
See Development Guide for detailed setup instructions.
# Run all tests
./tests/run-tests.sh
# Run specific test suite
dotnet test tests/SharpAudio.Api.Tests
dotnet test tests/SharpAudio.Api.IntegrationTests
# Opt-in: real-model circular TTS->ASR tests (downloads GB-scale weights, skipped by default and
# excluded from CI)
ASR_MODEL_TESTS=1 dotnet test tests/SharpAudio.Api.IntegrationTests --filter "Category=AsrModelTests"
# or
./tests/run-tests.sh asr-model-tests
SharpAudio is production-ready with:
See Deployment Guide for production deployment strategies.
Contributions are welcome! Please feel free to submit a Pull Request.
See LICENSE file for details.
MODEL_IDLE_TIMEOUT_SECONDS environment variable0 to disable automatic release and keep all loaded models in memoryUnit tests:
dotnet test tests/SharpAudio.Api.Tests/SharpAudio.Api.Tests.csproj
Integration tests with Docker:
# Run all tests (unit + integration)
./tests/run-tests.sh all
# Run specific test types
./tests/run-tests.sh unit
./tests/run-tests.sh integration
Or use Docker Compose directly:
docker compose -f docker-compose.test.yml up --abort-on-container-exit
GitHub Actions automatically runs tests on all pull requests:
The CI workflow:
main, master, or develop branchesSee .github/workflows/README.md for details.
HTTP Request → KokoroTtsSynthesizer
↓
KokoroTtsEngine
↓
┌───────────┴───────────┐
↓ ↓
EspeakWrapper OnnxRuntime
(phonemization) (inference)
↓ ↓
libespeak-ng.so.1 model.onnx
└───────────┬───────────┘
↓
PCM16 WAV Output
ASR follows an analogous path (AsrTranscriberRouter → WhisperAsrTranscriber/
NemotronAsrTranscriber → WhisperAsrEngine/NemotronAsrEngine), running in its own worker
process when SERVER_MODE enables ASR. See Architecture for the full
diagram covering both TTS and ASR.
C#
79.5%
Vue
13.3%
TypeScript
3.3%
Shell
1.8%
Dockerfile
1.0%
SharpAudio is a high-performance Text-to-Speech and Speech-to-Text REST API with OpenAI-compatible endpoints, highly vibe-coded:blush:. Built with .NET 10, it features a Vue.js frontend, Docker containerization, and support for multiple state-of-the-art TTS models (Kokoro, Supertonic-3) and ASR models (Whisper, Nemotron). Which task types a given deployment serves — TTS, ASR, or both — is controlled by a single SERVER_MODE environment variable.
/v1/audio/speech and /v1/audio/transcriptions endpointsSERVER_MODEdocker build --target# Clone the repository
git clone https://github.com/Fhrozen/SharpAudio.git
cd SharpAudio
# Start the service
docker compose up -d
# Access the web interface
open http://localhost:5768
# Or use the API directly
curl -X POST http://localhost:5768/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "kokoro-q4",
"input": "Hello world, this is a test of SharpAudio.",
"voice": "af_bella",
"speed": 1.0
}' \
--output speech.wav
The service will automatically download required models from Hugging Face on first startup.
By default the container only serves TTS. To enable Speech-to-Text as well, set SERVER_MODE:
# tts (default) | asr | both
SERVER_MODE=both docker compose up -d
# Transcribe an audio file (Whisper by default)
curl -X POST http://localhost:5768/v1/audio/transcriptions \
-F file=@sample.wav \
-F model=whisper-base
| Endpoint | Method | Description |
|---|---|---|
/health | GET | Health check endpoint |
/api/server-info | GET | Which task types (TTS/ASR) this server instance was started with |
/api/models | GET | List available TTS models with details (gated by SERVER_MODE) |
/api/asr-models | GET | List available ASR models with details (gated by SERVER_MODE) |
/v1/models | GET | OpenAI-compatible model listing (merges TTS + ASR catalogs) |
/v1/audio/speech | POST | OpenAI-compatible speech synthesis (gated by SERVER_MODE) |
/v1/audio/transcriptions | POST | OpenAI-compatible audio transcription (gated by SERVER_MODE) |
/swagger | GET | Interactive API documentation |
Text-to-Speech
| Model | Description | Languages | Voices | Quality |
|---|---|---|---|---|
kokoro-q4 | Quantized Kokoro 82M | 3+ | 5+ | Fast, low memory |
kokoro-full | Full precision Kokoro 82M | 3+ | 5+ | High quality |
supertonic-3 | Supertonic multilingual | 3+ | 10 styles | Production grade |
Speech-to-Text
| Model | Description | Languages | Notes |
|---|---|---|---|
whisper-base | OpenAI Whisper base (GGML, via whisper.cpp) | Auto-detect + multilingual | Whole-file batch transcription |
nemotron-3.5 | NVIDIA Nemotron 3.5 streaming ASR (FastConformer-RNNT, INT4 ONNX) | 35+ | Cache-aware chunked decoding |
See Models Documentation for detailed information about each model.
English (US/GB), Spanish, French, Hindi, Italian, Japanese, Portuguese (BR), Chinese (CN), Korean, German, Dutch, Arabic, Russian, Turkish, Polish, Swedish, Danish, Norwegian, Finnish, Greek, Czech, Romanian, Hungarian, Thai, Vietnamese, Indonesian, Hebrew, Ukrainian, and more.
Key environment variables:
HTTP_PORT=5768 # HTTP port
SERVER_MODE=tts # tts (default) | asr | both
MODEL_CACHE_DIR=/cache # Model cache directory
ESPEAK_DATA_DIR=/app/assets/espeak-ng-data # espeak-ng data
MODEL_IDLE_TIMEOUT_SECONDS=60 # Model unload timeout
See Configuration Guide for all options.
# Backend
cd src/SharpAudio.Api
dotnet restore
dotnet run
# Frontend
cd frontend
corepack enable pnpm
pnpm install
pnpm dev
See Development Guide for detailed setup instructions.
# Run all tests
./tests/run-tests.sh
# Run specific test suite
dotnet test tests/SharpAudio.Api.Tests
dotnet test tests/SharpAudio.Api.IntegrationTests
# Opt-in: real-model circular TTS->ASR tests (downloads GB-scale weights, skipped by default and
# excluded from CI)
ASR_MODEL_TESTS=1 dotnet test tests/SharpAudio.Api.IntegrationTests --filter "Category=AsrModelTests"
# or
./tests/run-tests.sh asr-model-tests
SharpAudio is production-ready with:
See Deployment Guide for production deployment strategies.
Contributions are welcome! Please feel free to submit a Pull Request.
See LICENSE file for details.
MODEL_IDLE_TIMEOUT_SECONDS environment variable0 to disable automatic release and keep all loaded models in memoryUnit tests:
dotnet test tests/SharpAudio.Api.Tests/SharpAudio.Api.Tests.csproj
Integration tests with Docker:
# Run all tests (unit + integration)
./tests/run-tests.sh all
# Run specific test types
./tests/run-tests.sh unit
./tests/run-tests.sh integration
Or use Docker Compose directly:
docker compose -f docker-compose.test.yml up --abort-on-container-exit
GitHub Actions automatically runs tests on all pull requests:
The CI workflow:
main, master, or develop branchesSee .github/workflows/README.md for details.
HTTP Request → KokoroTtsSynthesizer
↓
KokoroTtsEngine
↓
┌───────────┴───────────┐
↓ ↓
EspeakWrapper OnnxRuntime
(phonemization) (inference)
↓ ↓
libespeak-ng.so.1 model.onnx
└───────────┬───────────┘
↓
PCM16 WAV Output
ASR follows an analogous path (AsrTranscriberRouter → WhisperAsrTranscriber/
NemotronAsrTranscriber → WhisperAsrEngine/NemotronAsrEngine), running in its own worker
process when SERVER_MODE enables ASR. See Architecture for the full
diagram covering both TTS and ASR.
C#
79.5%
Vue
13.3%
TypeScript
3.3%
Shell
1.8%
Dockerfile
1.0%