Speech Recognition and Audio Models

11 repos

Libraries, models, and frameworks for automatic speech recognition (ASR) and audio processing, with a focus on PyTorch-based implementations. The cluster centers on end-to-end neural architectures like CTC, RNN-T, and TDT for converting speech to text, with several pre-trained Parakeet model variants at different scales. Developers here will find model implementations, training pipelines, and inference tools for building speech-to-text systems.

Python · 3
audio ·8,369
speech ·8,238
pytorch ·8,092
noise-suppression ·4,726
speech-enhancement ·4,726
deep-learning ·4,726
rust ·4,726
audio-processing ·2,939
io ·2,939
python ·2,939