An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
Python
4,541
309 commits
updated Aug 14, 2025
ClearerVoice-Studio is an open-source, AI-powered speech processing toolkit designed for researchers, developers, and end-users. It provides capabilities of speech enhancement, speech separation, speech super-resolution, target speaker extraction, and more. The toolkit provides state-of-the-art pre-trained models, along with training and inference scripts, all accessible from this repository.
Please leave your ⭐ on our GitHub to support this community project!
记得点击右上角的星星⭐来支持我们一下,您的支持是我们更新模型的最大动力!
demo_Numpy2Numpy.py.pip install clearvoice to use all the pretrained models in ClearVoice, see project description in PyPi link.This repository is organized into three main components: ClearVoice, Train, and SpeechScore.
ClearVoice offers a user-friendly solution for speech processing tasks such as speech denoising, separation, super-resolution, audio-visual target speaker extraction, and more. It is designed as a unified inference platform leveraged pre-trained models (e.g., FRCRN, MossFormer), all trained on extensive datasets. If you're looking for a tool to improve speech quality, ClearVoice is the perfect choice. Simply click on ClearVoice and follow our detailed instructions to get started.
For advanced researchers and developers, we provide model finetune and training scripts for all the tasks offerred in ClearVoice and more:
Contributors are welcomed to include more model architectures and tasks!
SpeechScore is a speech quality assessment toolkit. We include it here to evaluate different model performance. SpeechScore includes many popular speech metrics:
If you have any comments or questions about ClearerVoice-Studio, feel free to raise an issue in this repository or contact us directly at:
Alternatively, welcome to join our DingTalk group to share and discuss algorithms, technology, and user experience feedback. You may scan the following QR codes to join our official chat group.
Checkout some awesome Github repositories from Speech Lab of Institute for Intelligent Computing, Alibaba Group.
ClearerVoice-Studio contains third-party components and code modified from some open-source repos, including:
Speechbrain, ESPnet, TalkNet-ASD
104 followers · starred Dec 2024
218 followers · starred Dec 2024
401 followers · starred Dec 2024
553 followers · starred Dec 2024
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
Python
4,541
309 commits
updated Aug 14, 2025
ClearerVoice-Studio is an open-source, AI-powered speech processing toolkit designed for researchers, developers, and end-users. It provides capabilities of speech enhancement, speech separation, speech super-resolution, target speaker extraction, and more. The toolkit provides state-of-the-art pre-trained models, along with training and inference scripts, all accessible from this repository.
Please leave your ⭐ on our GitHub to support this community project!
记得点击右上角的星星⭐来支持我们一下,您的支持是我们更新模型的最大动力!
demo_Numpy2Numpy.py.pip install clearvoice to use all the pretrained models in ClearVoice, see project description in PyPi link.This repository is organized into three main components: ClearVoice, Train, and SpeechScore.
ClearVoice offers a user-friendly solution for speech processing tasks such as speech denoising, separation, super-resolution, audio-visual target speaker extraction, and more. It is designed as a unified inference platform leveraged pre-trained models (e.g., FRCRN, MossFormer), all trained on extensive datasets. If you're looking for a tool to improve speech quality, ClearVoice is the perfect choice. Simply click on ClearVoice and follow our detailed instructions to get started.
For advanced researchers and developers, we provide model finetune and training scripts for all the tasks offerred in ClearVoice and more:
Contributors are welcomed to include more model architectures and tasks!
SpeechScore is a speech quality assessment toolkit. We include it here to evaluate different model performance. SpeechScore includes many popular speech metrics:
If you have any comments or questions about ClearerVoice-Studio, feel free to raise an issue in this repository or contact us directly at:
Alternatively, welcome to join our DingTalk group to share and discuss algorithms, technology, and user experience feedback. You may scan the following QR codes to join our official chat group.
Checkout some awesome Github repositories from Speech Lab of Institute for Intelligent Computing, Alibaba Group.
ClearerVoice-Studio contains third-party components and code modified from some open-source repos, including:
Speechbrain, ESPnet, TalkNet-ASD
104 followers · starred Dec 2024
218 followers · starred Dec 2024
401 followers · starred Dec 2024
553 followers · starred Dec 2024