(This repository is in the experimental stage. The content may change without notice.)
Other languages
The structure of the decoder is designed with reference to FastSVC, StreamVC, Hifi-GAN, etc.
Low latency is achieved by using a "causal" convolution layer that does not refer to future information.
git clone https://github.com/uthree/fastersvc.git
pip3 install -r requirements.txt
The model pretrained with the JVS corpus is published here.
Train a model for basic voice conversion. At this stage, the model is not specialized for a specific speaker, but having a model that can perform basic voice synthesis allows for easy adaptation to a specific speaker with minimal adjustments.
Here are the steps:
python3 preprocess.py <Dataset directory>
python3 train_pe.py
python3 train_ce.py
sh
python3 train_dec.py
By adjusting the pre-trained model to a model specialized for conversion to a specific speaker, it is possible to create a more accurate model. This process takes much less time than pre-learning.
python3 preorpcess.py <Dataset directory>
python3 train_dec.py
python3 extract_index.py -o <Dictionary output destination (optional)>
-idx <dictionary file> option.-fp16 True to accelerate training with float16 if you have RTX series GPU.-b <number> to set batch size. default is 16.-e <number> to set epoch. default is 60.-d <device name> to set training device, default is cuda.inputsinputspython3 infer.py -t <target audio file>
-a <number from 0.0 to 1.0>.--normalize True.=d <device name>. Although it may not make much sense since it is originally high speed.-p <scale>. Useful for voice conversion between men and women.python3 audio_device_list.py
python3 infer_streaming.py -i <input device id> -o <output device id> -l <loopback device id> -t <target audio file>
(The loopback option is optional.)
This document is translated from Japanese using ChatGPT.
(This repository is in the experimental stage. The content may change without notice.)
Other languages
The structure of the decoder is designed with reference to FastSVC, StreamVC, Hifi-GAN, etc.
Low latency is achieved by using a "causal" convolution layer that does not refer to future information.
git clone https://github.com/uthree/fastersvc.git
pip3 install -r requirements.txt
The model pretrained with the JVS corpus is published here.
Train a model for basic voice conversion. At this stage, the model is not specialized for a specific speaker, but having a model that can perform basic voice synthesis allows for easy adaptation to a specific speaker with minimal adjustments.
Here are the steps:
python3 preprocess.py <Dataset directory>
python3 train_pe.py
python3 train_ce.py
sh
python3 train_dec.py
By adjusting the pre-trained model to a model specialized for conversion to a specific speaker, it is possible to create a more accurate model. This process takes much less time than pre-learning.
python3 preorpcess.py <Dataset directory>
python3 train_dec.py
python3 extract_index.py -o <Dictionary output destination (optional)>
-idx <dictionary file> option.-fp16 True to accelerate training with float16 if you have RTX series GPU.-b <number> to set batch size. default is 16.-e <number> to set epoch. default is 60.-d <device name> to set training device, default is cuda.inputsinputspython3 infer.py -t <target audio file>
-a <number from 0.0 to 1.0>.--normalize True.=d <device name>. Although it may not make much sense since it is originally high speed.-p <scale>. Useful for voice conversion between men and women.python3 audio_device_list.py
python3 infer_streaming.py -i <input device id> -o <output device id> -l <loopback device id> -t <target audio file>
(The loopback option is optional.)
This document is translated from Japanese using ChatGPT.