Pendrokar/xvapitch_nvidia

Model

xVASynth's xVAPitch (v3) type of voice models based on NVIDIA HIFI NeMo datasets.

8

16 commits

1 linked in READMEs

updated Apr 30, 2025

See the code

README

xVASynth's xVAPitch (v3) type of voice models based on NVIDIA HIFI NeMo datasets.

Models created by Dan Ruta, origin link:

Dataset supposed origin:

NameSynthesis Sample
ccby_nvidia_hifi_6671_MYour browser does not support the audio element.
ccby_nvidia_hifi_92_FYour browser does not support the audio element.
ccby_nvidia_hifi_6097_MYour browser does not support the audio element.
ccby_nv_hifi_11614_FYour browser does not support the audio element.
ccby_nvidia_hifi_11697_FYour browser does not support the audio element.
ccby_nvidia_hifi_12787_FYour browser does not support the audio element.
ccby_nvidia_hifi_6670_MYour browser does not support the audio element.
ccby_nvidia_hifi_8051_FYour browser does not support the audio element.
ccby_nvidia_hifi_9017_MYour browser does not support the audio element.
ccby_nvidia_hifi_9136_FYour browser does not support the audio element.

(These audio samples were created with the xVASynth Editor with the SR option (44kHz), not xVATrainer whose automatically created samples often sound different

Legal note: Although these datasets are licensed as CC BY 4.0, the base v3 model that these models are fine-tuned from, was pre-trained on non-permissive data.

v3 base model: https://huggingface.co/Pendrokar/xvapitch

audio
emotion
jp
text-to-speech

Contributors

Pendrokar

16 commits

Pendrokar/xvapitch_nvidia

Model

xVASynth's xVAPitch (v3) type of voice models based on NVIDIA HIFI NeMo datasets.

8

16 commits

1 linked in READMEs

updated Apr 30, 2025

See the code

README

xVASynth's xVAPitch (v3) type of voice models based on NVIDIA HIFI NeMo datasets.

Models created by Dan Ruta, origin link:

Dataset supposed origin:

NameSynthesis Sample
ccby_nvidia_hifi_6671_MYour browser does not support the audio element.
ccby_nvidia_hifi_92_FYour browser does not support the audio element.
ccby_nvidia_hifi_6097_MYour browser does not support the audio element.
ccby_nv_hifi_11614_FYour browser does not support the audio element.
ccby_nvidia_hifi_11697_FYour browser does not support the audio element.
ccby_nvidia_hifi_12787_FYour browser does not support the audio element.
ccby_nvidia_hifi_6670_MYour browser does not support the audio element.
ccby_nvidia_hifi_8051_FYour browser does not support the audio element.
ccby_nvidia_hifi_9017_MYour browser does not support the audio element.
ccby_nvidia_hifi_9136_FYour browser does not support the audio element.

(These audio samples were created with the xVASynth Editor with the SR option (44kHz), not xVATrainer whose automatically created samples often sound different

Legal note: Although these datasets are licensed as CC BY 4.0, the base v3 model that these models are fine-tuned from, was pre-trained on non-permissive data.

v3 base model: https://huggingface.co/Pendrokar/xvapitch

audio
emotion
jp
text-to-speech

Contributors

Pendrokar

16 commits