(This repository is experimental. Contents are subject to change without notice.)
(This file was generated by machine translation. may contain mistakes.)
Pretrained models are available here!

git clone https://github.com/uthree/tinyvc.git
pip3 install -r requirements.txt
Learn a model that performs basic speech conversion. At this stage, the model is not specialized for a specific speaker, but by preparing a model that can perform basic speech synthesis in advance, you can create a model that is specialized for a specific speaker with just a few adjustments. can be learned.
python3 preprocess.py <dataset directory>
python3 train_encoder.py
python3 train_decoder.py
By adjusting the pre-trained model to a model specialized for conversion to a specific speaker, it is possible to create a more accurate model. This process takes much less time than pre-learning.
python3 preprocess.py
python3 train_decoder.py
python3 extract_index.py -o <Dictionary output destination (optional)>
-idx <dictionary file> option.
The default dictionary file output destination is models/index.pt.-fp16 True allows learning using 16-bit floating point numbers. Possible only for RTX series GPUs.-b <number>. Default is 16.-e <number>. Default is 60.-d <device name>. Default is cuda.inputs folder.inputs folderpython3 infer.py -t <target audio file>
Also, if you use a dictionary file,
python3 infer.py -idx <dictionary file>
-d <device name>. Although it may not make much sense since it is originally high speed.-p <scale>. Useful for voice conversion between men and women. 12 raises it one octave.--no-chunking True to incerease quality, but this mode requires more RAM.python3 audio_device_list.py
python3 infer_streaming.py -i <input device ID> -o <output device ID> -l <loopback device ID> -t <target audio file>
(It works even without the loopback option.)
(This repository is experimental. Contents are subject to change without notice.)
(This file was generated by machine translation. may contain mistakes.)
Pretrained models are available here!

git clone https://github.com/uthree/tinyvc.git
pip3 install -r requirements.txt
Learn a model that performs basic speech conversion. At this stage, the model is not specialized for a specific speaker, but by preparing a model that can perform basic speech synthesis in advance, you can create a model that is specialized for a specific speaker with just a few adjustments. can be learned.
python3 preprocess.py <dataset directory>
python3 train_encoder.py
python3 train_decoder.py
By adjusting the pre-trained model to a model specialized for conversion to a specific speaker, it is possible to create a more accurate model. This process takes much less time than pre-learning.
python3 preprocess.py
python3 train_decoder.py
python3 extract_index.py -o <Dictionary output destination (optional)>
-idx <dictionary file> option.
The default dictionary file output destination is models/index.pt.-fp16 True allows learning using 16-bit floating point numbers. Possible only for RTX series GPUs.-b <number>. Default is 16.-e <number>. Default is 60.-d <device name>. Default is cuda.inputs folder.inputs folderpython3 infer.py -t <target audio file>
Also, if you use a dictionary file,
python3 infer.py -idx <dictionary file>
-d <device name>. Although it may not make much sense since it is originally high speed.-p <scale>. Useful for voice conversion between men and women. 12 raises it one octave.--no-chunking True to incerease quality, but this mode requires more RAM.python3 audio_device_list.py
python3 infer_streaming.py -i <input device ID> -o <output device ID> -l <loopback device ID> -t <target audio file>
(It works even without the loopback option.)