hugging-apps/irodori-tts-anime-demo

Space

19

stars

4

commits

2

linked in READMEs

Sep 6, 2026

updated

gradio
mcp-server

README

Irodori-TTS-v4.1-Anime

Demo for phasefield-audio/Irodori-TTS-v4.1-Anime, a ~766M-parameter Japanese text-to-speech model fine-tuned from Aratako/Irodori-TTS-v4.1-Small on anime-style speech data.

The architecture is a Rectified Flow Diffusion Transformer (RF-DiT) that generates continuous Semantic-DACVAE-Japanese-32dim latents, decoded to 48 kHz audio.

Features

  • Text → speech in Japanese, with automatic duration prediction.
  • Style captions — describe the voice/emotion in Japanese and the model follows it. This fine-tune was annotated independently of the base model, so caption behaviour differs from Irodori-TTS-v4.1-Small.
  • Zero-shot voice cloning — upload one or more reference clips (concatenated, up to 120 s).
  • Emoji style controls — insert control emoji (whisper, sigh, giggle, …) inline in the text.
  • Generated audio is watermarked with SilentCipher.

Credits

Model by phasefield-audio. Inference code (irodori_tts/) vendored from Aratako/Irodori-TTS (MIT).

License

MIT, subject to the base model's ethical restrictions regarding impersonation and misinformation.

Contributors

multimodalart

4 commits

hugging-apps/irodori-tts-anime-demo

Space

19

stars

4

commits

2

linked in READMEs

Sep 6, 2026

updated

gradio
mcp-server

README

Irodori-TTS-v4.1-Anime

Demo for phasefield-audio/Irodori-TTS-v4.1-Anime, a ~766M-parameter Japanese text-to-speech model fine-tuned from Aratako/Irodori-TTS-v4.1-Small on anime-style speech data.

The architecture is a Rectified Flow Diffusion Transformer (RF-DiT) that generates continuous Semantic-DACVAE-Japanese-32dim latents, decoded to 48 kHz audio.

Features

  • Text → speech in Japanese, with automatic duration prediction.
  • Style captions — describe the voice/emotion in Japanese and the model follows it. This fine-tune was annotated independently of the base model, so caption behaviour differs from Irodori-TTS-v4.1-Small.
  • Zero-shot voice cloning — upload one or more reference clips (concatenated, up to 120 s).
  • Emoji style controls — insert control emoji (whisper, sigh, giggle, …) inline in the text.
  • Generated audio is watermarked with SilentCipher.

Credits

Model by phasefield-audio. Inference code (irodori_tts/) vendored from Aratako/Irodori-TTS (MIT).

License

MIT, subject to the base model's ethical restrictions regarding impersonation and misinformation.

Contributors

multimodalart

4 commits