A Japanese text-to-speech model fine-tuned from Aratako/Irodori-TTS-v4.1-Small using anime-style speech data.
The base model's annotation pipeline is not publicly documented, so the fine-tuning data was annotated independently. Consequently, caption conditioning and emoji controls may behave differently from the base model.
The full-precision checkpoint is available at the repository root.
Quantized variants are provided in the following subdirectories:
int8-weight-onlyint8-dynamicint4-weight-onlyfloat8-weight-onlyfloat8-dynamicFor inference and installation instructions, see the original Irodori-TTS repository.
This model follows the same MIT License and ethical restrictions as the base model.
9 commits
A Japanese text-to-speech model fine-tuned from Aratako/Irodori-TTS-v4.1-Small using anime-style speech data.
The base model's annotation pipeline is not publicly documented, so the fine-tuning data was annotated independently. Consequently, caption conditioning and emoji controls may behave differently from the base model.
The full-precision checkpoint is available at the repository root.
Quantized variants are provided in the following subdirectories:
int8-weight-onlyint8-dynamicint4-weight-onlyfloat8-weight-onlyfloat8-dynamicFor inference and installation instructions, see the original Irodori-TTS repository.
This model follows the same MIT License and ethical restrictions as the base model.
9 commits