Speech Recognition and Speaker Diarization

12 repos

Audio processing models and systems for automatic speech recognition (ASR), speaker identification, and speaker diarization—determining who spoke when in multi-speaker environments. The cluster includes streaming and real-time variants (sortformer, parakeet), speaker verification models (titanet), and dictionary-of-context resources (DiCoW), all built primarily in PyTorch. This is an applied machine learning area focusing on end-to-end speech understanding pipelines.

audio ·906
speech ·788
pytorch ·721
speaker-recognition ·589
automatic-speech-recognition ·566
speaker-diarization ·563
nemo ·551
NeMo ·551
model-index ·551
Conformer ·424