Speech processing and audio analysis

30 repos

Libraries, models, and tools for analyzing, segmenting, and understanding speech and speaker characteristics in audio. The cluster centers on speaker diarization (identifying who spoke when), voice activity detection, overlapped speech handling, and speaker segmentation—core tasks in speech understanding pipelines. Most repos build on or integrate with pyannote-audio, a widely-used framework for speaker-related audio analysis tasks, alongside complementary work in audio codecs and speech representation learning.

audio ·11,067
speech ·11,067
voice ·9,470
pyannote ·9,469
pyannote-audio ·9,469
speaker ·9,438
voice-activity-detection ·9,169
overlapped-speech-detection ·8,961
speaker-change-detection ·8,202
speaker-diarization ·8,202