Speech Recognition and Audio Processing

33 repos

Libraries, models, and tools for automatic speech recognition (ASR), speaker diarization, and audio processing tasks. The cluster spans multiple deployment contexts—from on-device inference using CoreML and ONNX Runtime to browser-based implementations with transformers.js—making these resources useful for building speech-enabled applications across platforms. Central repositories include optimized model variants like GPA and specialized audio models such as LFM2.5 and Raon-SpeechChat, reflecting a focus on making speech AI practical and accessible across different hardware and runtime constraints.

Python · 6
C · 2
Swift · 1
audio ·14,539
speech ·9,829
automatic-speech-recognition ·7,332
python ·6,256
vad ·5,189
voice-activity-detection ·5,043
real-time ·4,976
speech-to-text ·4,891
speaker-diarization ·4,741
pytorch ·4,129