Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
See the code
OpenSpeech provides reference implementations of various ASR modeling papers and three languages recipe to perform tasks on automatic speech recognition. We aim to make ASR technology easier to use for everyone.
OpenSpeech is backed by the two powerful libraries — PyTorch-Lightning and Hydra.
Various features are available in the above two libraries, including Multi-GPU and TPU training, Mixed-precision, and hierarchical configuration management.
We appreciate any kind of feedback or contribution. Feel free to proceed with small issues like bug fixes, documentation improvement. For major contributions and new features, please discuss with the collaborators in corresponding issues.
OpenSpeech is a framework for making end-to-end speech recognizers. End-to-end (E2E) automatic speech recognition (ASR) is an emerging paradigm in the field of neural network-based speech recognition that offers multiple benefits. Traditional “hybrid” ASR systems, which are comprised of an acoustic model, language model, and pronunciation model, require separate training of these components, each of which can be complex.
For example, training of an acoustic model is a multi-stage process of model training and time alignment between the speech acoustic feature sequence and output label sequence. In contrast, E2E ASR is a single integrated approach with a much simpler training pipeline with models that operate at low audio frame rates. This reduces the training time, decoding time, and allows joint optimization with downstream processing such as natural language understanding.
Because of these advantages, many end-to-end speech recognition related open sources have emerged. But, Many of them are based on basic PyTorch or Tensorflow, it is very difficult to use various functions such as mixed-precision, multi-node training, and TPU training etc. However, with frameworks such as PyTorch-Lighting, these features can be easily used. So we have created a speech recognition framework that introduced PyTorch-Lightning and Hydra for easy use of these advanced features.
pl.LightingDataModule and Tokenizer classes.We support all the models below. Note that, the important concepts of the model have been implemented to match, but the details of the implementation may vary.
We use Hydra to control all the training configurations. If you are not familiar with Hydra we recommend visiting the Hydra website. Generally, Hydra is an open-source framework that simplifies the development of research applications by providing the ability to create a hierarchical configuration dynamically. If you want to know how we used Hydra, we recommend you to read here.
We support LibriSpeech, KsponSpeech, and AISHELL-1.
LibriSpeech is a corpus of approximately 1,000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data was derived from reading audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Aishell is an open-source Chinese Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. 400 people from different accent areas in China were invited to participate in the recording, which was conducted in a quiet indoor environment using high fidelity microphone and downsampled to 16kHz.
KsponSpeech is a large-scale spontaneous speech corpus of Korean. This corpus contains 969 hours of general open-domain dialog utterances, spoken by about 2,000 native Korean speakers in a clean environment. All data were constructed by recording the dialogue of two people freely conversing on a variety of topics and manually transcribing the utterances. To start training, the KsponSpeech dataset must be prepared in advance. To download KsponSpeech, you need permission from AI Hub.
LibriSpeech/test-other/8188/269288/8188-269288-0052.flac ▁ANNIE ' S ▁MANNER ▁WAS ▁VERY ▁MYSTERIOUS 4039 20 5 531 17 84 2352
LibriSpeech/test-other/8188/269288/8188-269288-0053.flac ▁ANNIE ▁DID ▁NOT ▁MEAN ▁TO ▁CONFIDE ▁IN ▁ANYONE ▁THAT ▁NIGHT ▁AND ▁THE ▁KIND EST ▁THING ▁WAS ▁TO ▁LEAVE ▁HER ▁A LONE 4039 99 35 251 9 4758 11 2454 16 199 6 4 323 200 255 17 9 370 30 10 492
LibriSpeech/test-other/8188/269288/8188-269288-0054.flac ▁TIRED ▁OUT ▁LESLIE ▁HER SELF ▁DROPP ED ▁A SLEEP 1493 70 4708 30 115 1231 7 10 1706
LibriSpeech/test-other/8188/269288/8188-269288-0055.flac ▁ANNIE ▁IS ▁THAT ▁YOU ▁SHE ▁CALL ED ▁OUT 4039 34 16 25 37 208 7 70
LibriSpeech/test-other/8188/269288/8188-269288-0056.flac ▁THERE ▁WAS ▁NO ▁REPLY ▁BUT ▁THE ▁SOUND ▁OF ▁HURRY ING ▁STEPS ▁CAME ▁QUICK ER ▁AND ▁QUICK ER ▁NOW ▁AND ▁THEN ▁THEY ▁WERE ▁INTERRUPTED ▁BY ▁A ▁GROAN 57 17 56 1368 33 4 489 8 1783 14 1381 133 571 49 6 571 49 82 6 76 45 54 2351 44 10 3154
LibriSpeech/test-other/8188/269288/8188-269288-0057.flac ▁OH ▁THIS ▁WILL ▁KILL ▁ME ▁MY ▁HEART ▁WILL ▁BREAK ▁THIS ▁WILL ▁KILL ▁ME 299 46 71 669 50 41 235 71 977 46 71 669 50
...
...
You can simply train with LibriSpeech dataset like below:
conformer-lstm model with filter-bank features on GPU.$ python3 ./openspeech_cli/hydra_train.py \
dataset=librispeech \
dataset.dataset_download=True \
dataset.dataset_path=$DATASET_PATH \
dataset.manifest_file_path=$MANIFEST_FILE_PATH \
tokenizer=libri_subword \
model=conformer_lstm \
audio=fbank \
lr_scheduler=warmup_reduce_lr_on_plateau \
trainer=gpu \
criterion=cross_entropy
You can simply train with KsponSpeech dataset like below:
listen-attend-spell model with mel-spectrogram features On TPU:$ python3 ./openspeech_cli/hydra_train.py \
dataset=ksponspeech \
dataset.dataset_path=$DATASET_PATH \
dataset.manifest_file_path=$MANIFEST_FILE_PATH \
dataset.test_dataset_path=$TEST_DATASET_PATH \
dataset.test_manifest_dir=$TEST_MANIFEST_DIR \
tokenizer=kspon_character \
model=listen_attend_spell \
audio=melspectrogram \
lr_scheduler=warmup_reduce_lr_on_plateau \
trainer=tpu \
criterion=cross_entropy
You can simply train with AISHELL-1 dataset like below:
quartznet model with mfcc features On GPU with FP16:$ python3 ./openspeech_cli/hydra_train.py \
dataset=aishell \
dataset.dataset_path=$DATASET_PATH \
dataset.dataset_download=True \
dataset.manifest_file_path=$MANIFEST_FILE_PATH \
tokenizer=aishell_character \
model=quartznet15x5 \
audio=mfcc \
lr_scheduler=warmup_reduce_lr_on_plateau \
trainer=gpu-fp16 \
criterion=ctc
listen_attend_spell model:$ python3 ./openspeech_cli/hydra_eval.py \
audio=melspectrogram \
eval.dataset_path=$DATASET_PATH \
eval.checkpoint_path=$CHECKPOINT_PATH \
eval.manifest_file_path=$MANIFEST_FILE_PATH \
model=listen_attend_spell \
tokenizer=kspon_character \
tokenizer.vocab_path=$VOCAB_FILE_PATH \
listen_attend_spell, conformer_lstm models with ensemble:$ python3 ./openspeech_cli/hydra_eval.py \
audio=melspectrogram \
eval.model_names=(listen_attend_spell, conformer_lstm) \
eval.dataset_path=$DATASET_PATH \
eval.checkpoint_paths=($CHECKPOINT_PATH1, $CHECKPOINT_PATH2) \
eval.ensemble_weights=(0.3, 0.7) \
eval.ensemble_method=weighted \
eval.manifest_file_path=$MANIFEST_FILE_PATH
dataset.dataset_path: $BASE_PATH/KsponSpeech$BASE_PATH/KsponSpeech
├── KsponSpeech_01
├── KsponSpeech_02
├── KsponSpeech_03
├── KsponSpeech_04
└── KsponSpeech_05
dataset.test_dataset_path: $BASE_PATH/KsponSpeech_eval$BASE_PATH/KsponSpeech_eval
├── eval_clean
└── eval_other
dataset.test_manifest_dir: $BASE_PATH/KsponSpeech_scripts$BASE_PATH/KsponSpeech_scripts
├── eval_clean.trn
└── eval_other.trn
Language model training requires only data to be prepared in the following format:
openspeech is a framework for making end-to-end speech recognizers.
end to end automatic speech recognition is an emerging paradigm in the field of neural network-based speech recognition that offers multiple benefits.
because of these advantages, many end-to-end speech recognition related open sources have emerged.
...
...
Note that you need to use the same vocabulary as the acoustic model.
lstm_lm model:$ python3 ./openspeech_cli/hydra_lm_train.py \
dataset=lm \
dataset.dataset_path=../../../lm.txt \
tokenizer=kspon_character \
tokenizer.vocab_path=../../../labels.csv \
model=lstm_lm \
lr_scheduler=tri_stage \
trainer=gpu \
criterion=perplexity
This project recommends Python 3.7 or higher. We recommend creating a new virtual environment for this project (using virtual env or conda).
pip install numpy (Refer here for problem installing Numpy).conda install -c conda-forge librosa (Refer here for problem installing librosa)pip install torchaudio==0.6.0 (Refer here for problem installing torchaudio)pip install sentencepiece (Refer here for problem installing sentencepiece)pip install pytorch-lightning (Refer here for problem installing pytorch-lightning)pip install hydra-core --upgrade (Refer here for problem installing hydra)You can install OpenSpeech with pypi.
pip install openspeech-core
Currently we only support installation from source code using setuptools. Checkout the source code and run the following commands:
$ ./install.sh
For faster training install NVIDIA's apex library:
$ git clone https://github.com/NVIDIA/apex
$ cd apex
# ------------------------
# OPTIONAL: on your cluster you might need to load CUDA 10 or 9
# depending on how you installed PyTorch
# see available modules
module avail
# load correct CUDA before install
module load cuda-10.0
# ------------------------
# make sure you've loaded a cuda version > 4.0 and < 7.0
module load gcc-6.1.0
$ pip install -v --no-cache-dir --global-option="--cpp_ext" --global-option="--cuda_ext" ./
If you have any questions, bug reports, and feature requests, please open an issue on Github.
We appreciate any kind of feedback or contribution. Feel free to proceed with small issues like bug fixes, documentation improvement. For major contributions and new features, please discuss with the collaborators in corresponding issues.
We follow PEP-8 for code style. Especially the style of docstrings is important to generate documentation.
This project is licensed under the MIT LICENSE - see the LICENSE.md file for details
If you use the system for academic work, please cite:
@GITHUB{2021-OpenSpeech,
author = {Kim, Soohwan and Ha, Sangchun and Cho, Soyoung},
author email = {sh951011@gmail.com, seomk9896@gmail.com, soyoung.cho@kaist.ac.kr}
title = {OpenSpeech: Open-Source Toolkit for End-to-End Speech Recognition},
howpublished = {\url{https://github.com/openspeech-team/openspeech}},
docs = {\url{https://openspeech-team.github.io/openspeech}},
year = {2021}
}
Python
99.8%
Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
See the code
OpenSpeech provides reference implementations of various ASR modeling papers and three languages recipe to perform tasks on automatic speech recognition. We aim to make ASR technology easier to use for everyone.
OpenSpeech is backed by the two powerful libraries — PyTorch-Lightning and Hydra.
Various features are available in the above two libraries, including Multi-GPU and TPU training, Mixed-precision, and hierarchical configuration management.
We appreciate any kind of feedback or contribution. Feel free to proceed with small issues like bug fixes, documentation improvement. For major contributions and new features, please discuss with the collaborators in corresponding issues.
OpenSpeech is a framework for making end-to-end speech recognizers. End-to-end (E2E) automatic speech recognition (ASR) is an emerging paradigm in the field of neural network-based speech recognition that offers multiple benefits. Traditional “hybrid” ASR systems, which are comprised of an acoustic model, language model, and pronunciation model, require separate training of these components, each of which can be complex.
For example, training of an acoustic model is a multi-stage process of model training and time alignment between the speech acoustic feature sequence and output label sequence. In contrast, E2E ASR is a single integrated approach with a much simpler training pipeline with models that operate at low audio frame rates. This reduces the training time, decoding time, and allows joint optimization with downstream processing such as natural language understanding.
Because of these advantages, many end-to-end speech recognition related open sources have emerged. But, Many of them are based on basic PyTorch or Tensorflow, it is very difficult to use various functions such as mixed-precision, multi-node training, and TPU training etc. However, with frameworks such as PyTorch-Lighting, these features can be easily used. So we have created a speech recognition framework that introduced PyTorch-Lightning and Hydra for easy use of these advanced features.
pl.LightingDataModule and Tokenizer classes.We support all the models below. Note that, the important concepts of the model have been implemented to match, but the details of the implementation may vary.
We use Hydra to control all the training configurations. If you are not familiar with Hydra we recommend visiting the Hydra website. Generally, Hydra is an open-source framework that simplifies the development of research applications by providing the ability to create a hierarchical configuration dynamically. If you want to know how we used Hydra, we recommend you to read here.
We support LibriSpeech, KsponSpeech, and AISHELL-1.
LibriSpeech is a corpus of approximately 1,000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data was derived from reading audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Aishell is an open-source Chinese Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. 400 people from different accent areas in China were invited to participate in the recording, which was conducted in a quiet indoor environment using high fidelity microphone and downsampled to 16kHz.
KsponSpeech is a large-scale spontaneous speech corpus of Korean. This corpus contains 969 hours of general open-domain dialog utterances, spoken by about 2,000 native Korean speakers in a clean environment. All data were constructed by recording the dialogue of two people freely conversing on a variety of topics and manually transcribing the utterances. To start training, the KsponSpeech dataset must be prepared in advance. To download KsponSpeech, you need permission from AI Hub.
LibriSpeech/test-other/8188/269288/8188-269288-0052.flac ▁ANNIE ' S ▁MANNER ▁WAS ▁VERY ▁MYSTERIOUS 4039 20 5 531 17 84 2352
LibriSpeech/test-other/8188/269288/8188-269288-0053.flac ▁ANNIE ▁DID ▁NOT ▁MEAN ▁TO ▁CONFIDE ▁IN ▁ANYONE ▁THAT ▁NIGHT ▁AND ▁THE ▁KIND EST ▁THING ▁WAS ▁TO ▁LEAVE ▁HER ▁A LONE 4039 99 35 251 9 4758 11 2454 16 199 6 4 323 200 255 17 9 370 30 10 492
LibriSpeech/test-other/8188/269288/8188-269288-0054.flac ▁TIRED ▁OUT ▁LESLIE ▁HER SELF ▁DROPP ED ▁A SLEEP 1493 70 4708 30 115 1231 7 10 1706
LibriSpeech/test-other/8188/269288/8188-269288-0055.flac ▁ANNIE ▁IS ▁THAT ▁YOU ▁SHE ▁CALL ED ▁OUT 4039 34 16 25 37 208 7 70
LibriSpeech/test-other/8188/269288/8188-269288-0056.flac ▁THERE ▁WAS ▁NO ▁REPLY ▁BUT ▁THE ▁SOUND ▁OF ▁HURRY ING ▁STEPS ▁CAME ▁QUICK ER ▁AND ▁QUICK ER ▁NOW ▁AND ▁THEN ▁THEY ▁WERE ▁INTERRUPTED ▁BY ▁A ▁GROAN 57 17 56 1368 33 4 489 8 1783 14 1381 133 571 49 6 571 49 82 6 76 45 54 2351 44 10 3154
LibriSpeech/test-other/8188/269288/8188-269288-0057.flac ▁OH ▁THIS ▁WILL ▁KILL ▁ME ▁MY ▁HEART ▁WILL ▁BREAK ▁THIS ▁WILL ▁KILL ▁ME 299 46 71 669 50 41 235 71 977 46 71 669 50
...
...
You can simply train with LibriSpeech dataset like below:
conformer-lstm model with filter-bank features on GPU.$ python3 ./openspeech_cli/hydra_train.py \
dataset=librispeech \
dataset.dataset_download=True \
dataset.dataset_path=$DATASET_PATH \
dataset.manifest_file_path=$MANIFEST_FILE_PATH \
tokenizer=libri_subword \
model=conformer_lstm \
audio=fbank \
lr_scheduler=warmup_reduce_lr_on_plateau \
trainer=gpu \
criterion=cross_entropy
You can simply train with KsponSpeech dataset like below:
listen-attend-spell model with mel-spectrogram features On TPU:$ python3 ./openspeech_cli/hydra_train.py \
dataset=ksponspeech \
dataset.dataset_path=$DATASET_PATH \
dataset.manifest_file_path=$MANIFEST_FILE_PATH \
dataset.test_dataset_path=$TEST_DATASET_PATH \
dataset.test_manifest_dir=$TEST_MANIFEST_DIR \
tokenizer=kspon_character \
model=listen_attend_spell \
audio=melspectrogram \
lr_scheduler=warmup_reduce_lr_on_plateau \
trainer=tpu \
criterion=cross_entropy
You can simply train with AISHELL-1 dataset like below:
quartznet model with mfcc features On GPU with FP16:$ python3 ./openspeech_cli/hydra_train.py \
dataset=aishell \
dataset.dataset_path=$DATASET_PATH \
dataset.dataset_download=True \
dataset.manifest_file_path=$MANIFEST_FILE_PATH \
tokenizer=aishell_character \
model=quartznet15x5 \
audio=mfcc \
lr_scheduler=warmup_reduce_lr_on_plateau \
trainer=gpu-fp16 \
criterion=ctc
listen_attend_spell model:$ python3 ./openspeech_cli/hydra_eval.py \
audio=melspectrogram \
eval.dataset_path=$DATASET_PATH \
eval.checkpoint_path=$CHECKPOINT_PATH \
eval.manifest_file_path=$MANIFEST_FILE_PATH \
model=listen_attend_spell \
tokenizer=kspon_character \
tokenizer.vocab_path=$VOCAB_FILE_PATH \
listen_attend_spell, conformer_lstm models with ensemble:$ python3 ./openspeech_cli/hydra_eval.py \
audio=melspectrogram \
eval.model_names=(listen_attend_spell, conformer_lstm) \
eval.dataset_path=$DATASET_PATH \
eval.checkpoint_paths=($CHECKPOINT_PATH1, $CHECKPOINT_PATH2) \
eval.ensemble_weights=(0.3, 0.7) \
eval.ensemble_method=weighted \
eval.manifest_file_path=$MANIFEST_FILE_PATH
dataset.dataset_path: $BASE_PATH/KsponSpeech$BASE_PATH/KsponSpeech
├── KsponSpeech_01
├── KsponSpeech_02
├── KsponSpeech_03
├── KsponSpeech_04
└── KsponSpeech_05
dataset.test_dataset_path: $BASE_PATH/KsponSpeech_eval$BASE_PATH/KsponSpeech_eval
├── eval_clean
└── eval_other
dataset.test_manifest_dir: $BASE_PATH/KsponSpeech_scripts$BASE_PATH/KsponSpeech_scripts
├── eval_clean.trn
└── eval_other.trn
Language model training requires only data to be prepared in the following format:
openspeech is a framework for making end-to-end speech recognizers.
end to end automatic speech recognition is an emerging paradigm in the field of neural network-based speech recognition that offers multiple benefits.
because of these advantages, many end-to-end speech recognition related open sources have emerged.
...
...
Note that you need to use the same vocabulary as the acoustic model.
lstm_lm model:$ python3 ./openspeech_cli/hydra_lm_train.py \
dataset=lm \
dataset.dataset_path=../../../lm.txt \
tokenizer=kspon_character \
tokenizer.vocab_path=../../../labels.csv \
model=lstm_lm \
lr_scheduler=tri_stage \
trainer=gpu \
criterion=perplexity
This project recommends Python 3.7 or higher. We recommend creating a new virtual environment for this project (using virtual env or conda).
pip install numpy (Refer here for problem installing Numpy).conda install -c conda-forge librosa (Refer here for problem installing librosa)pip install torchaudio==0.6.0 (Refer here for problem installing torchaudio)pip install sentencepiece (Refer here for problem installing sentencepiece)pip install pytorch-lightning (Refer here for problem installing pytorch-lightning)pip install hydra-core --upgrade (Refer here for problem installing hydra)You can install OpenSpeech with pypi.
pip install openspeech-core
Currently we only support installation from source code using setuptools. Checkout the source code and run the following commands:
$ ./install.sh
For faster training install NVIDIA's apex library:
$ git clone https://github.com/NVIDIA/apex
$ cd apex
# ------------------------
# OPTIONAL: on your cluster you might need to load CUDA 10 or 9
# depending on how you installed PyTorch
# see available modules
module avail
# load correct CUDA before install
module load cuda-10.0
# ------------------------
# make sure you've loaded a cuda version > 4.0 and < 7.0
module load gcc-6.1.0
$ pip install -v --no-cache-dir --global-option="--cpp_ext" --global-option="--cuda_ext" ./
If you have any questions, bug reports, and feature requests, please open an issue on Github.
We appreciate any kind of feedback or contribution. Feel free to proceed with small issues like bug fixes, documentation improvement. For major contributions and new features, please discuss with the collaborators in corresponding issues.
We follow PEP-8 for code style. Especially the style of docstrings is important to generate documentation.
This project is licensed under the MIT LICENSE - see the LICENSE.md file for details
If you use the system for academic work, please cite:
@GITHUB{2021-OpenSpeech,
author = {Kim, Soohwan and Ha, Sangchun and Cho, Soyoung},
author email = {sh951011@gmail.com, seomk9896@gmail.com, soyoung.cho@kaist.ac.kr}
title = {OpenSpeech: Open-Source Toolkit for End-to-End Speech Recognition},
howpublished = {\url{https://github.com/openspeech-team/openspeech}},
docs = {\url{https://openspeech-team.github.io/openspeech}},
year = {2021}
}
Python
99.8%