MioTTS-0.6B is a lightweight, high-speed Text-to-Speech (TTS) model based on an LLM architecture. It is designed to generate high-quality speech in English and Japanese while maintaining low latency and minimal resource usage.
This model supports zero-shot voice cloning and is built on top of the efficient neural audio codec MioCodec-25Hz-24kHz.
We offer a range of model sizes to suit different performance and resource requirements.
| Model Name | Parameters | Base Model | License | RTF (Real-Time Factor) |
|---|---|---|---|---|
| MioTTS-0.1B | 0.1B | tiiuae/Falcon-H1-Tiny-Multilingual-100M-Base | Falcon-LLM License | 0.04 - 0.05 |
| MioTTS-0.4B | 0.4B | LiquidAI/LFM2-350M | LFM Open License v1.0 | 0.035 - 0.045 |
| MioTTS-0.6B | 0.6B | Qwen/Qwen3-0.6B-Base | Apache 2.0 | 0.055 - 0.065 |
| MioTTS-1.2B | 1.2B | LiquidAI/LFM2.5-1.2B-Base | LFM Open License v1.0 | 0.065 - 0.075 |
| MioTTS-1.7B | 1.7B | Qwen/Qwen3-1.7B-Base | Apache 2.0 | 0.10 - 0.11 |
| MioTTS-2.6B | 2.6B | LiquidAI/LFM2-2.6B | LFM Open License v1.0 | 0.135 - 0.145 |
RTF values represent the range observed when generating approximately 15 seconds of audio across multiple runs. Measured on an NVIDIA RTX 5090 using vLLM 0.15.1.
We provide a dedicated repository for inference, including installation instructions and example WebUI.
๐ GitHub: Aratako/MioTTS-Inference
Below are some samples generated by MioTTS-0.6B.
Note: The reference audio samples below were generated using Aratako/T5Gemma-TTS-2b-2b and gemini-2.5-pro-tts.
| Case | Text | Reference Audio | Generated Audio |
|---|---|---|---|
| English 1 | "The old library was silent, save for the gentle ticking of a clock somewhere in the shadows. As I ran my fingers along the dusty spines of the books, I felt a strange sense of nostalgia, as if I had lived a thousand lives within these walls." | ||
| English 2 | "Hey! I haven't seen you in ages. Do you want to grab some coffee later? I've got so much to tell you!" | ||
| Japanese 1 | "ๆฐ่ฑกๅบใซใใใพใใจใๅคงๅใฎๅฐ้ขจ10ๅทใฏใๆๆฅใฎๆใๆนใซใใใฆ้ขๆฑๅฐๆนใซๆฅ่ฟใใ่ฆ่พผใฟใงใใๆฒฟๅฒธ้จใงใฏ้ซๆณขใซ่ญฆๆใๅฟ ่ฆใงใใ" | ||
| Japanese 2 | "ใใฎๆฃฎใซใฏใๅคใ่จใไผใใใใใพใใใๆใๆใ้ซใๆใๅคใ้ใใซ่ณใๆพใพใใฐใ้ขจใฎๆญๅฃฐใ่ใใใใจใใใฎใงใใ็งใฏๅไฟกๅ็ใงใใใใใใฎๅคใ็ขบใใซ่ชฐใใ็งใๅผใถๅฃฐใ่ใใใฎใงใใ" |
This model is released under the Apache 2.0.
While this model is released under a permissive license, we aim to promote responsible AI development and urge users to respect the rights of others.
If you use MioTTS in your research or project, please cite it as follows:
@misc{miotts,
author = {Chihiro Arata},
title = {MioTTS: Lightweight and Fast LLM-based Text-to-Speech},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face repository},
howpublished = {\url{https://huggingface.co/collections/Aratako/miotts}}
}
5 commits
MioTTS-0.6B is a lightweight, high-speed Text-to-Speech (TTS) model based on an LLM architecture. It is designed to generate high-quality speech in English and Japanese while maintaining low latency and minimal resource usage.
This model supports zero-shot voice cloning and is built on top of the efficient neural audio codec MioCodec-25Hz-24kHz.
We offer a range of model sizes to suit different performance and resource requirements.
| Model Name | Parameters | Base Model | License | RTF (Real-Time Factor) |
|---|---|---|---|---|
| MioTTS-0.1B | 0.1B | tiiuae/Falcon-H1-Tiny-Multilingual-100M-Base | Falcon-LLM License | 0.04 - 0.05 |
| MioTTS-0.4B | 0.4B | LiquidAI/LFM2-350M | LFM Open License v1.0 | 0.035 - 0.045 |
| MioTTS-0.6B | 0.6B | Qwen/Qwen3-0.6B-Base | Apache 2.0 | 0.055 - 0.065 |
| MioTTS-1.2B | 1.2B | LiquidAI/LFM2.5-1.2B-Base | LFM Open License v1.0 | 0.065 - 0.075 |
| MioTTS-1.7B | 1.7B | Qwen/Qwen3-1.7B-Base | Apache 2.0 | 0.10 - 0.11 |
| MioTTS-2.6B | 2.6B | LiquidAI/LFM2-2.6B | LFM Open License v1.0 | 0.135 - 0.145 |
RTF values represent the range observed when generating approximately 15 seconds of audio across multiple runs. Measured on an NVIDIA RTX 5090 using vLLM 0.15.1.
We provide a dedicated repository for inference, including installation instructions and example WebUI.
๐ GitHub: Aratako/MioTTS-Inference
Below are some samples generated by MioTTS-0.6B.
Note: The reference audio samples below were generated using Aratako/T5Gemma-TTS-2b-2b and gemini-2.5-pro-tts.
| Case | Text | Reference Audio | Generated Audio |
|---|---|---|---|
| English 1 | "The old library was silent, save for the gentle ticking of a clock somewhere in the shadows. As I ran my fingers along the dusty spines of the books, I felt a strange sense of nostalgia, as if I had lived a thousand lives within these walls." | ||
| English 2 | "Hey! I haven't seen you in ages. Do you want to grab some coffee later? I've got so much to tell you!" | ||
| Japanese 1 | "ๆฐ่ฑกๅบใซใใใพใใจใๅคงๅใฎๅฐ้ขจ10ๅทใฏใๆๆฅใฎๆใๆนใซใใใฆ้ขๆฑๅฐๆนใซๆฅ่ฟใใ่ฆ่พผใฟใงใใๆฒฟๅฒธ้จใงใฏ้ซๆณขใซ่ญฆๆใๅฟ ่ฆใงใใ" | ||
| Japanese 2 | "ใใฎๆฃฎใซใฏใๅคใ่จใไผใใใใใพใใใๆใๆใ้ซใๆใๅคใ้ใใซ่ณใๆพใพใใฐใ้ขจใฎๆญๅฃฐใ่ใใใใจใใใฎใงใใ็งใฏๅไฟกๅ็ใงใใใใใใฎๅคใ็ขบใใซ่ชฐใใ็งใๅผใถๅฃฐใ่ใใใฎใงใใ" |
This model is released under the Apache 2.0.
While this model is released under a permissive license, we aim to promote responsible AI development and urge users to respect the rights of others.
If you use MioTTS in your research or project, please cite it as follows:
@misc{miotts,
author = {Chihiro Arata},
title = {MioTTS: Lightweight and Fast LLM-based Text-to-Speech},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face repository},
howpublished = {\url{https://huggingface.co/collections/Aratako/miotts}}
}
5 commits