Text-to-Speech and Audio Generation

29 repos

Tools and models for synthesizing speech from text and converting between audio modalities, with a focus on efficient implementations across edge and mobile platforms. The cluster includes neural TTS engines, GGUF-quantized model variants for resource-constrained inference, and cross-platform implementations—particularly in Swift for iOS/macOS deployment. Central repos like VoxCPM, Soprano, and Kokoro represent compact language models optimized for real-time speech synthesis.

Swift · 3
C++ · 1
Rust · 1
text-to-speech ·1,792
tts ·1,202
mlx ·1,148
mlx-audio ·793
stt ·774
speech-to-text ·774
mlx-audio-swift ·774
mlx-swift-audio ·774
speech-to-speech ·774
en ·621