Vietnamese zero-shot TTS / voice cloning fine-tuned from ZipVoice.
checkpoint-1860000.pt, FP16 inference state dict7000 total hours, including roughly 6500 hours of Vietnamese and 500 hours of EnglishSimpleTokenizer, character-level, 244 tokens24 kHzcharactr/vocos-mel-24khzThe wrapper loads the largest checkpoint-<step>.pt automatically and uses soe-vinorm for Vietnamese text normalization.
Generated with checkpoint-1860000.pt, the current wrapper flow, and the demo text in demo/demo_text.txt.
Đinh-Quyết
Nhã-Uyên
MC
git clone https://github.com/iamdinhthuan/ViZipvoice.git
cd ViZipvoice
pip install -r requirements.txt
export PYTHONPATH="$PWD:$PYTHONPATH"
python3 -m zipvoice.bin.infer_vizipvoice \
--prompt-wav prompt.wav \
--prompt-text "Xin chào, đây là giọng mẫu của tôi." \
--text "ViZipVoice có thể tổng hợp giọng nói tiếng Việt từ một đoạn mẫu ngắn." \
--res-wav-path output.wav
The CLI downloads this model repo by default. Use --model-dir models/ViZipvoice after downloading files locally.
from zipvoice.vizipvoice import ViZipVoiceTTS
tts = ViZipVoiceTTS()
metrics = tts.synthesize(
prompt_wav="prompt.wav",
prompt_text="Xin chào, đây là giọng mẫu của tôi.",
text="Đây là câu tiếng Việt được sinh bởi ViZipVoice.",
output_path="output.wav",
)
print(metrics)
audio/ contains 30 reference prompts. Each audio file has a sidecar .txt transcript with the same basename:
audio/Đinh-Quyết.mp3
audio/Đinh-Quyết.txt
Names only keep the audio/person name; the original lar_* prefix and Pro suffix are removed. The Gradio app reads this sidecar format automatically.
huggingface-cli download contextboxai/ViZipvoice \
--local-dir models/ViZipvoice \
--local-dir-use-symlinks False
python3 egs/zipvoice/gradio_app.py --exp-dir models/ViZipvoice
The CLI, Python wrapper, and Gradio app use the same default flow:
soe-vinorm, then clean spaces around punctuation;1-word sentence: use at least 24 steps and speed=0.6;2-4 word sentence: use speed=0.8;Useful knobs:
--no-vietnamese-normalize
--no-split-sentences
--crossfade-ms 80
--silence-ms 180
--fade-in-ms 20
--fade-out-ms 80
checkpoint-1860000.pt: latest FP16 checkpointconfig.json, model.json: model configtokens.txt: Vietnamese character tokenizeraudio/: 30 reference audios plus .txt transcriptsdemo/: regenerated audio demos and metadata.jsonvizipvoice.py: wrapper mirrored from GitHubThis model can clone voices from short audio prompts. Use only voices you own or have explicit permission to use. Do not use it for impersonation, fraud, harassment, misinformation, or other harmful content.
Apache License 2.0. Please also credit the original ZipVoice project.
Vietnamese zero-shot TTS / voice cloning fine-tuned from ZipVoice.
checkpoint-1860000.pt, FP16 inference state dict7000 total hours, including roughly 6500 hours of Vietnamese and 500 hours of EnglishSimpleTokenizer, character-level, 244 tokens24 kHzcharactr/vocos-mel-24khzThe wrapper loads the largest checkpoint-<step>.pt automatically and uses soe-vinorm for Vietnamese text normalization.
Generated with checkpoint-1860000.pt, the current wrapper flow, and the demo text in demo/demo_text.txt.
Đinh-Quyết
Nhã-Uyên
MC
git clone https://github.com/iamdinhthuan/ViZipvoice.git
cd ViZipvoice
pip install -r requirements.txt
export PYTHONPATH="$PWD:$PYTHONPATH"
python3 -m zipvoice.bin.infer_vizipvoice \
--prompt-wav prompt.wav \
--prompt-text "Xin chào, đây là giọng mẫu của tôi." \
--text "ViZipVoice có thể tổng hợp giọng nói tiếng Việt từ một đoạn mẫu ngắn." \
--res-wav-path output.wav
The CLI downloads this model repo by default. Use --model-dir models/ViZipvoice after downloading files locally.
from zipvoice.vizipvoice import ViZipVoiceTTS
tts = ViZipVoiceTTS()
metrics = tts.synthesize(
prompt_wav="prompt.wav",
prompt_text="Xin chào, đây là giọng mẫu của tôi.",
text="Đây là câu tiếng Việt được sinh bởi ViZipVoice.",
output_path="output.wav",
)
print(metrics)
audio/ contains 30 reference prompts. Each audio file has a sidecar .txt transcript with the same basename:
audio/Đinh-Quyết.mp3
audio/Đinh-Quyết.txt
Names only keep the audio/person name; the original lar_* prefix and Pro suffix are removed. The Gradio app reads this sidecar format automatically.
huggingface-cli download contextboxai/ViZipvoice \
--local-dir models/ViZipvoice \
--local-dir-use-symlinks False
python3 egs/zipvoice/gradio_app.py --exp-dir models/ViZipvoice
The CLI, Python wrapper, and Gradio app use the same default flow:
soe-vinorm, then clean spaces around punctuation;1-word sentence: use at least 24 steps and speed=0.6;2-4 word sentence: use speed=0.8;Useful knobs:
--no-vietnamese-normalize
--no-split-sentences
--crossfade-ms 80
--silence-ms 180
--fade-in-ms 20
--fade-out-ms 80
checkpoint-1860000.pt: latest FP16 checkpointconfig.json, model.json: model configtokens.txt: Vietnamese character tokenizeraudio/: 30 reference audios plus .txt transcriptsdemo/: regenerated audio demos and metadata.jsonvizipvoice.py: wrapper mirrored from GitHubThis model can clone voices from short audio prompts. Use only voices you own or have explicit permission to use. Do not use it for impersonation, fraud, harassment, misinformation, or other harmful content.
Apache License 2.0. Please also credit the original ZipVoice project.