Speech Recognition & Audio Models

169 repos across 9 sub-areas

Automatic speech recognition (ASR) systems and multilingual audio processing models, primarily built on wav2vec2 and Conformer architectures. The cluster centers on pre-trained models fine-tuned for speech-to-text tasks across dozens of languages, along with supporting libraries for audio feature extraction, voice activity detection, and acoustic modeling. Most repos are model checkpoints and inference utilities rather than foundational frameworks, reflecting a focus on applied speech technology.

Cluster 637677

39 repos

Multilingual Speech Recognition Models

34 repos

Pre-trained multilingual automatic speech recognition models built on wav2vec 2.0 and transformer architectures, enabling speech-to-text across diverse languages including Hungarian, Finnish, Persian, Chinese, Arabic, and Greek. The cluster primarily consists of fine-tuned model repositories that apply the XLSR-53 (cross-lingual speech representations) framework to language-specific datasets, leveraging PyTorch and the Hugging Face transformers library for production speech recognition applications.

Cluster 637676

31 repos

Cluster 637672

29 repos

Cluster 637671

14 repos

Speech Recognition and Audio Models

11 repos

Libraries, models, and frameworks for automatic speech recognition (ASR) and audio processing, with a focus on PyTorch-based implementations. The cluster centers on end-to-end neural architectures like CTC, RNN-T, and TDT for converting speech to text, with several pre-trained Parakeet model variants at different scales. Developers here will find model implementations, training pipelines, and inference tools for building speech-to-text systems.

Speech Processing and Audio ML

7 repos

Deep learning approaches to speech enhancement, synthesis, and real-time audio processing. The cluster centers on neural codec models and end-to-end speech systems built with PyTorch, including tools for training speech enhancement networks, full-duplex audio interaction, and audio compression. Repositories span from low-level audio DSP to high-level conversational AI applications, with a focus on practical speech quality improvements and efficient model deployment.

Cluster 637674

3 repos

Cluster 637670

1 repos