12 repos
Audio processing models and systems for automatic speech recognition (ASR), speaker identification, and speaker diarization—determining who spoke when in multi-speaker environments. The cluster includes streaming and real-time variants (sortformer, parakeet), speaker verification models (titanet), and dictionary-of-context resources (DiCoW), all built primarily in PyTorch. This is an applied machine learning area focusing on end-to-end speech understanding pipelines.