A high-speed, high-fidelity Text-to-Speech model optimized for real-time conversational AI applications.
KaniTTS uses a two-stage pipeline combining a large language model with an efficient audio codec for exceptional speed and audio quality. The architecture generates compressed token representations through a backbone LLM, then rapidly synthesizes waveforms via neural audio codec, achieving extremely low latency.
Key Specifications:
Nvidia RTX 5080 Benchmarks:
Pretraining:
| Text | Audio |
|---|---|
| I do believe Marsellus Wallace, MY husband, YOUR boss, told you to take me out and do WHATEVER I WANTED. | |
| What do we say to the god of death? Not today! | |
| What do you call a lawyer with an IQ of 60? Your honor | |
| You mean, let me understand this cause, you know maybe it's me, it's a little fucked up maybe, but I'm funny how, I mean funny like I'm a clown, I amuse you? |
Models:
Examples:
Links:
Built on top of LiquidAI LFM2 350M as the backbone and Nvidia NanoCodec for audio processing.
Prohibited activities include:
By using this model, you agree to comply with these restrictions and all applicable laws.
Have a question, feedback, or need support? Please fill out our contact form and we'll get back to you as soon as possible.
A high-speed, high-fidelity Text-to-Speech model optimized for real-time conversational AI applications.
KaniTTS uses a two-stage pipeline combining a large language model with an efficient audio codec for exceptional speed and audio quality. The architecture generates compressed token representations through a backbone LLM, then rapidly synthesizes waveforms via neural audio codec, achieving extremely low latency.
Key Specifications:
Nvidia RTX 5080 Benchmarks:
Pretraining:
| Text | Audio |
|---|---|
| I do believe Marsellus Wallace, MY husband, YOUR boss, told you to take me out and do WHATEVER I WANTED. | |
| What do we say to the god of death? Not today! | |
| What do you call a lawyer with an IQ of 60? Your honor | |
| You mean, let me understand this cause, you know maybe it's me, it's a little fucked up maybe, but I'm funny how, I mean funny like I'm a clown, I amuse you? |
Models:
Examples:
Links:
Built on top of LiquidAI LFM2 350M as the backbone and Nvidia NanoCodec for audio processing.
Prohibited activities include:
By using this model, you agree to comply with these restrictions and all applicable laws.
Have a question, feedback, or need support? Please fill out our contact form and we'll get back to you as soon as possible.