Text-to-Speech (TTS) Systems

37 repos

Neural text-to-speech models and implementations for converting written text into spoken audio. The cluster centers on deep learning TTS architectures, audio generation pipelines, and model checkpoints (including the Irodori 500M and DOTS SOAR variants), with emphasis on practical implementations using modern formats like safetensors for efficient model distribution. Repositories here cover both model training frameworks and inference tools for building voice synthesis applications.

text-to-speech ·2,234
tts ·1,948
safetensors ·1,828
audio ·1,207
en ·999
voice ·822
speech-synthesis ·742
speech ·613
llama ·544
moshi ·379