30 repos
Libraries, models, and tools for analyzing, segmenting, and understanding speech and speaker characteristics in audio. The cluster centers on speaker diarization (identifying who spoke when), voice activity detection, overlapped speech handling, and speaker segmentation—core tasks in speech understanding pipelines. Most repos build on or integrate with pyannote-audio, a widely-used framework for speaker-related audio analysis tasks, alongside complementary work in audio codecs and speech representation learning.