Irodori-TTS-500M is a Japanese Text-to-Speech model based on a Rectified Flow Diffusion Transformer (RF-DiT) architecture. The architecture and training design largely follow Echo-TTS, using DACVAE continuous latents as the generation target. It supports zero-shot voice cloning from reference audio.
A unique feature of this model is emoji-based style and sound effect control โ by inserting specific emojis into the input text, you can control speaking styles, emotions, and even sound effects in the generated audio.
EMOJI_ANNOTATIONS.md for the full list of supported emojis and their effects.The model (approximately 500M parameters) consists of three main components:
Audio is represented as continuous latent sequences via the DACVAE codec (128-dim), enabling high-quality 48kHz waveform reconstruction.
Basic Japanese text-to-speech generation (without reference audio).
| Case | Text | Generated Audio |
|---|---|---|
| Sample 1 | "ใ้ป่ฉฑใใใใจใใใใใพใใใใ ใใพ้ป่ฉฑใๅคงๅคๆททใฟๅใฃใฆใใใพใใๆใๅ ฅใใพใใใ็บไฟก้ณใฎใใจใซใใ็จไปถใใ่ฉฑใใใ ใใใ" | |
| Sample 2 | "ใใฎๆฃฎใซใฏใๅคใ่จใไผใใใใใพใใใๆใๆใ้ซใๆใๅคใ้ใใซ่ณใๆพใพใใฐใ้ขจใฎๆญๅฃฐใ่ใใใใจใใใฎใงใใ็งใฏๅไฟกๅ็ใงใใใใใใฎๅคใ็ขบใใซ่ชฐใใ็งใๅผใถๅฃฐใ่ใใใฎใงใใ" |
Examples of controlling speaking style and effects with emojis. For the full list of supported emojis, see EMOJI_ANNOTATIONS.md.
| Case | Text (with Emoji) | Generated Audio |
|---|---|---|
| Sample 1 | ใชใผใซใใฉใใใใฎ๏ผโฆใ๏ผใใฃใจ่ฟใฅใใฆใปใใ๏ผโฆ๐๐ฎโ๐จ๐๐ฎโ๐จใใใใใฎใๅฅฝใใชใใ ๏ผ | |
| Sample 2 | ใใ โฆ๐ญใใใชใซ้ ทใใใจใ่จใใชใใงโฆ๐ญ | |
| Sample 3 | ๐คง๐คงใใใใญใ้ขจ้ชๅผใใกใใฃใฆใฆ๐คงโฆๅคงไธๅคซใใใ ใฎ้ขจ้ชใ ใใใใๆฒปใใ๐ฅบ |
Examples of cloning a voice from a reference audio clip.
| Case | Reference Audio | Generated Audio |
|---|---|---|
| Example 1 | ||
| Example 2 |
For inference code, installation instructions, and training scripts, please refer to the GitHub repository:
๐ GitHub: Aratako/Irodori-TTS
The model was trained on a high-quality Japanese speech dataset. To enable the emoji-based style control, the training texts were enriched with emoji annotations. These annotations were automatically generated and labeled using a fine-tuned model based on Qwen/Qwen3-Omni-30B-A3B-Instruct.
This model is released under MIT.
In addition to the license terms, the following ethical restrictions apply:
This project builds upon the following works:
We would also like to extend our special thanks to Respair for the inspiration behind the emoji annotation feature.
If you use Irodori-TTS in your research or project, please cite it as follows:
@misc{irodori-tts,
author = {Chihiro Arata},
title = {Irodori-TTS: A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face repository},
howpublished = {\url{https://huggingface.co/Aratako/Irodori-TTS-500M}}
}
3 commits
Irodori-TTS-500M is a Japanese Text-to-Speech model based on a Rectified Flow Diffusion Transformer (RF-DiT) architecture. The architecture and training design largely follow Echo-TTS, using DACVAE continuous latents as the generation target. It supports zero-shot voice cloning from reference audio.
A unique feature of this model is emoji-based style and sound effect control โ by inserting specific emojis into the input text, you can control speaking styles, emotions, and even sound effects in the generated audio.
EMOJI_ANNOTATIONS.md for the full list of supported emojis and their effects.The model (approximately 500M parameters) consists of three main components:
Audio is represented as continuous latent sequences via the DACVAE codec (128-dim), enabling high-quality 48kHz waveform reconstruction.
Basic Japanese text-to-speech generation (without reference audio).
| Case | Text | Generated Audio |
|---|---|---|
| Sample 1 | "ใ้ป่ฉฑใใใใจใใใใใพใใใใ ใใพ้ป่ฉฑใๅคงๅคๆททใฟๅใฃใฆใใใพใใๆใๅ ฅใใพใใใ็บไฟก้ณใฎใใจใซใใ็จไปถใใ่ฉฑใใใ ใใใ" | |
| Sample 2 | "ใใฎๆฃฎใซใฏใๅคใ่จใไผใใใใใพใใใๆใๆใ้ซใๆใๅคใ้ใใซ่ณใๆพใพใใฐใ้ขจใฎๆญๅฃฐใ่ใใใใจใใใฎใงใใ็งใฏๅไฟกๅ็ใงใใใใใใฎๅคใ็ขบใใซ่ชฐใใ็งใๅผใถๅฃฐใ่ใใใฎใงใใ" |
Examples of controlling speaking style and effects with emojis. For the full list of supported emojis, see EMOJI_ANNOTATIONS.md.
| Case | Text (with Emoji) | Generated Audio |
|---|---|---|
| Sample 1 | ใชใผใซใใฉใใใใฎ๏ผโฆใ๏ผใใฃใจ่ฟใฅใใฆใปใใ๏ผโฆ๐๐ฎโ๐จ๐๐ฎโ๐จใใใใใฎใๅฅฝใใชใใ ๏ผ | |
| Sample 2 | ใใ โฆ๐ญใใใชใซ้ ทใใใจใ่จใใชใใงโฆ๐ญ | |
| Sample 3 | ๐คง๐คงใใใใญใ้ขจ้ชๅผใใกใใฃใฆใฆ๐คงโฆๅคงไธๅคซใใใ ใฎ้ขจ้ชใ ใใใใๆฒปใใ๐ฅบ |
Examples of cloning a voice from a reference audio clip.
| Case | Reference Audio | Generated Audio |
|---|---|---|
| Example 1 | ||
| Example 2 |
For inference code, installation instructions, and training scripts, please refer to the GitHub repository:
๐ GitHub: Aratako/Irodori-TTS
The model was trained on a high-quality Japanese speech dataset. To enable the emoji-based style control, the training texts were enriched with emoji annotations. These annotations were automatically generated and labeled using a fine-tuned model based on Qwen/Qwen3-Omni-30B-A3B-Instruct.
This model is released under MIT.
In addition to the license terms, the following ethical restrictions apply:
This project builds upon the following works:
We would also like to extend our special thanks to Respair for the inspiration behind the emoji annotation feature.
If you use Irodori-TTS in your research or project, please cite it as follows:
@misc{irodori-tts,
author = {Chihiro Arata},
title = {Irodori-TTS: A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face repository},
howpublished = {\url{https://huggingface.co/Aratako/Irodori-TTS-500M}}
}
3 commits