Speech Processing and Audio ML

7 repos

Deep learning approaches to speech enhancement, synthesis, and real-time audio processing. The cluster centers on neural codec models and end-to-end speech systems built with PyTorch, including tools for training speech enhancement networks, full-duplex audio interaction, and audio compression. Repositories span from low-level audio DSP to high-level conversational AI applications, with a focus on practical speech quality improvements and efficient model deployment.

speech ·221
audio ·221
conversational ·81
full-duplex ·81
english ·68
text-to-speech ·51
anime ·48
galgame ·48
japanese ·48
text ·48