141 repos across 6 sub-areas
Cluster 462266
39 repos
Speech Recognition and Audio Processing
33 repos
Libraries, models, and tools for automatic speech recognition (ASR), speaker diarization, and audio processing tasks. The cluster spans multiple deployment contexts—from on-device inference using CoreML and ONNX Runtime to browser-based implementations with transformers.js—making these resources useful for building speech-enabled applications across platforms. Central repositories include optimized model variants like GPA and specialized audio models such as LFM2.5 and Raon-SpeechChat, reflecting a focus on making speech AI practical and accessible across different hardware and runtime constraints.
Speech processing and audio analysis
30 repos
Libraries, models, and tools for analyzing, segmenting, and understanding speech and speaker characteristics in audio. The cluster centers on speaker diarization (identifying who spoke when), voice activity detection, overlapped speech handling, and speaker segmentation—core tasks in speech understanding pipelines. Most repos build on or integrate with pyannote-audio, a widely-used framework for speaker-related audio analysis tasks, alongside complementary work in audio codecs and speech representation learning.
Multilingual Speech Recognition Models
21 repos
Pre-trained multilingual automatic speech recognition models built on wav2vec 2.0 and transformer architectures, enabling speech-to-text across diverse languages including Hungarian, Finnish, Persian, Chinese, Arabic, and Greek. The cluster primarily consists of fine-tuned model repositories that apply the XLSR-53 (cross-lingual speech representations) framework to language-specific datasets, leveraging PyTorch and the Hugging Face transformers library for production speech recognition applications.
Cluster 462267
12 repos
Cluster 462270
6 repos