🎤 A curated list of the latest and most influential tools, models, and resources in the Text-to-Speech sector. 🌟 Star if you like it! 🌟
194
119 commits
updated Sep 14, 2026
Welcome to the most comprehensive, meticulously curated, and continuously updated list of Text-to-Speech (TTS) resources. Whether you are looking for the best open-source TTS models of 2026, searching for low-latency TTS APIs for AI agents, or exploring high-fidelity voice cloning for content creation, you've found the right place.
[!TIP] Looking for the best ElevenLabs alternatives? This repository tracks the rapidly evolving landscape of both commercial SaaS and local-first neural speech synthesis.
Text-to-Speech technology has moved beyond robotic voices. Today, it powers:
The landscape of AI voice synthesis has shifted from basic concatenation to advanced Generative Speech Models. Key highlights:
Leading platforms offering robust, scalable, and high-quality Text-to-Speech APIs and services for various applications. Sorted by Company Scale / Valuation (Descending).
| 🛠️ Service/Model | 🏢 Organization | 🌟 Key Features | 💰 Min. Monthly Subscription | 👥 Company Size | 🔗 Link |
|---|---|---|---|---|---|
| NVIDIA NeMo | NVIDIA | Platform for building, training, and deploying generative AI models, including TTS and ASR. | Free (API Credits) / Enterprise | $3.3T+ (Market Cap) | NVIDIA NeMo |
| Azure AI Speech | Microsoft | High-quality neural voices with advanced fine-tuning, emotion, and enterprise scalability. | Free (0.5M chars/mo) / PAYG | $3.2T+ (Market Cap) | Azure AI Speech |
| Google Cloud TTS | Powerful TTS API with a large variety of natural-sounding voices and extensive customization. | Free (1M+ chars/mo) / PAYG | $2.2T+ (Market Cap) | Google Cloud TTS | |
| AWS Polly | Amazon | Generative, Neural and Standard TTS voices with deep AWS ecosystem integration. | Free (1M+ chars/mo) / PAYG | $1.9T+ (Market Cap) | AWS Polly |
| OpenAI TTS | OpenAI | High-quality, real-time streaming TTS models for applications requiring natural AI voices. | Pay-as-you-go ($5 free credit) | $852B (Valuation) | OpenAI TTS |
| ElevenLabs | ElevenLabs | State-of-the-art AI voice generator offering realistic voices, voice cloning, and AI dubbing. | Free (10k chars/mo) / $5 | $11B (Valuation) | ElevenLabs |
| Speechify | Speechify | Highly popular consumer text-to-speech and developer Voice API with natural and premium voices. | Free (Basic Voices) / $139/yr | $1.5B (Valuation) | Speechify |
| Deepgram Aura | Deepgram | Specializing in low-latency TTS designed for real-time conversational AI and virtual interactions. | Free ($200 credit) / PAYG | $1.2B (Valuation) | Deepgram Aura |
| Inworld AI | Inworld AI | Character-driven conversational voice engine and real-time speech generation for interactive agents. | Free (Basic / 100k API credits) / $20 | $500M (Valuation) | Inworld AI |
| Hume AI (EVI) | Hume AI | Empathic Voice Interface with emotional prosody detection and low-latency expressive speech. | Free ($10 credit) / PAYG ($0.036/min) | $250M (Valuation) | Hume AI |
| Cartesia Sonic | Cartesia | Sub-100ms ultra-low latency TTS designed for real-time AI agents. | Free (API Credits) / PAYG | $200M (Valuation) | Cartesia |
| Gradium | Gradium | Real-time TTS and STT for voice agents. 158ms P50 time-to-first-audio, streaming and instant voice cloning. | Free (45k credits) / $13 | $100M (Seed Raised) | Gradium |
| Murf.ai | Murf.ai | AI voiceovers with a built-in video editor, ideal for creators and presentations. | Free (10 mins total) / $19 | $46M (Valuation) | Murf.ai |
| WellSaid Labs | WellSaid Labs | Enterprise AI voice platform with natural studio voices, fine phonetic control, and brand voice avatars. | Free (7-day trial / 50 clips) / $49 | $40M (Valuation) | WellSaid Labs |
| Resemble AI | Resemble AI | High-fidelity voice cloning, neural watermarking, deepfake detection, and real-time speech synthesis. | Free (Free trial) / $29 | $30M (Valuation) | Resemble AI |
| LMNT | LMNT | Lightning-fast TTS API with exceptional naturalness, great for interactive voice apps. | Free (API Credits) / PAYG | $21M (Estimated) | LMNT |
| Lovo.ai (Genny) | LOVO | Next-gen AI voice generator with 500+ voices in 100+ languages and granular pitch/emphasis controls. | Free (14-day Pro trial) / $29 | $20M (Estimated) | Lovo.ai |
| Play.ht | Play.ht | Professional AI voices and "Ultra-Realistic" studio editor for long-form content. | Free (12.5k chars) / $39 | $15M (Valuation) | Play.ht |
| Smallest.ai (Waves) | Smallest.ai | Lightning-fast TTS API with sub-100ms latency designed for real-time conversational agents and voice bots. | Free (10k chars/mo) / PAYG ($0.08/1k chars) | $12M (Valuation) | Smallest.ai |
| Rime Labs | Rime | Ultra-fast expressive voice API built specifically for interactive conversational AI applications. | Free ($5 credit) / PAYG | $10M (Estimated) | Rime |
| Soniox TTS | Soniox | Real-time streaming TTS API for conversational AI voice agents in 60+ languages with multilingual voices. | Pay-as-you-go (~$0.70/hr) | $10M (Estimated) | Soniox |
| Neets.ai | Neets | Extremely fast and affordable TTS APIs starting at $0.0004 per 1k characters. | Free (API Credits) / PAYG | $5M (Estimated) | Neets.ai |
| Gandr | Gandr | TTS API for voice agents. Word error rate 1.98 percent against a 2.17 percent human reference on the same scorer, one voice in 23 languages, every render watermarked. Python/JS SDKs, LiveKit plugin, MCP server. | From $10/mo. Flat-rate unmetered streams from $150/mo | <$1M (Indie) | Gandr |
| Spokio | Spokio | Offline macOS text-to-speech app with local voice cloning, batch export, and no cloud uploads. | Free (API Credits) / Enterprise | <$1M (Indie) | Spokio |
| PHANTOM VOICES | PHANTOM VOICES | 10 free professional AI voice clones via public REST API. Zero cost, commercial rights cleared. 29 platform configs (Vapi, Retell AI, etc). Multilingual (9+ languages). AI-powered recommendation. | Free (API Credits) / Enterprise | <$1M (Indie) | PHANTOM VOICES |
| RunAPI ElevenLabs SDK | RunAPI | Multi-language SDKs for ElevenLabs text-to-speech, dialogue generation, sound effects, transcription, and audio isolation workflows. | Pay-as-you-go | <$1M (Indie) | RunAPI ElevenLabs SDK |
| Audexum | Audexum | Text-to-speech and speech-to-text in one API: 43 voices, 32 TTS languages, 25 STT languages. ElevenLabs-compatible endpoint, so switching is a base-URL change. EU-hosted, every output watermarked. | Free (30k credits at signup, 3k/mo after) / EUR 4 | <$1M (Indie) | Audexum |
If you are looking for free text-to-speech models for commercial use or want to run TTS locally on a CPU, these open-source projects provide the best balance of quality and privacy. Sorted by GitHub Star Counts (Descending).
| 🛠️ Service/Model | 🏢 Organization | 🌟 Key Features | 🗣️ Primary Language | 📁 Github_Repository |
|---|---|---|---|---|
| 🐸 Coqui TTS | Coqui (community) | Supports 1100+ languages, zero-shot voice cloning, and fine-tuning. Note: Coqui AI (the company) shut down in late 2023; actively maintained by the community at idiap/coqui-ai-TTS. | Python / Multilingual | |
| GPT-SoVITS | RVC-Boss | Zero-shot & few-shot voice cloning. Requires only 5 seconds of sample audio for cross-lingual synthesis. | Python / PyTorch | |
| Bark | Suno | Transformer-based text-to-audio model capable of highly expressive speech, music, laughs, and sighs. | Python / PyTorch | |
| ChatTTS | 2noise | Conversational text-to-speech model specially optimized for dialogue and natural conversational flow. | Python | |
| OpenVoice | MyShell | Highly versatile and instant voice cloning that requires only a short audio clip. | Python | |
| Fish Speech | Fish Audio | SOTA multilingual, multi-speaker model with superior naturalness. | Python | |
| Chatterbox | Resemble AI | Advanced neural voice synthesis with emotion control and high-fidelity cloning. | Python | |
| CosyVoice | Alibaba | Excellent multilingual and zero-shot voice cloning model capable of high fidelity. | Python | |
| KittenTTS | KittenML | ONNX-based library for low-latency TTS without requiring a GPU. | Python / ONNX | |
| F5-TTS | SWivid | Flow Matching TTS. Incredible naturalness and prosody using DiT architectures. | Python | |
| Tortoise-TTS | James Betker | Powerful multi-voice TTS system known for its exceptional voice cloning capabilities. | Python | |
| VALL-E-X | Plachtaa | Open-source implementation of Microsoft's VALL-E X for zero-shot cross-lingual voice cloning. | Python / PyTorch | |
| Piper | Rhasspy | Fastest local TTS. Optimized for low-end hardware and offline use. | C++ / Python | |
| Amphion | Amphion | Open-source audio, music and speech generation toolkit containing multiple SOTA TTS models. | Python | |
| VoiceCraft | Jason Li et al. | Token-infilling neural codec model for zero-shot speech editing and TTS synthesis. | Python / PyTorch | |
| Kokoro-82M | Hexgrad | Best SOTA CPU TTS. Ultra-fast, studio quality, only 82M parameters. | ONNX / Python | |
| StyleTTS 2 | yl4579 | Human-level TTS. Uses style diffusion, adversarial training, and large SLMs without phoneme duration models. | Python / PyTorch | |
| MeloTTS | MyShell | Ultra-fast multilingual TTS running smoothly on CPU across English, Spanish, French, Chinese, Japanese, Korean. | Python / PyTorch | |
| sherpa-onnx | Next-gen Kaldi | Offline multi-platform speech synthesis engine supporting VITS, Piper, and Kokoro on embedded/mobile/desktop. | C++ / Python / Go | |
| Parler-TTS | Hugging Face | Lightweight, controllable speech generation with high naturalness. | Python | HF |
| GLM-4-Voice | Zhipu AI | End-to-end voice model supporting real-time speech generation, emotion alteration, and bilingual dialogue. | Python / PyTorch | |
| Matcha-TTS | Shivam Mehta | Fast TTS architecture employing conditional flow matching, producing highly natural output. | Python | |
| LocalMode | LocalMode | In-browser TTS. Runs Kokoro (29 voices) and other AI models 100% in the browser via WebGPU/WASM. No server, no API keys, offline after first load. | JavaScript / TypeScript | |
| Vocello | PowerBeef | Native Mac & iPhone app. Qwen3-TTS with preset speakers, natural-language voice design, and voice cloning. Runs entirely on Apple Silicon with no Python runtime, faster than realtime on an 8 GB M2. | Swift / MLX | |
| loudkit | LoudReader | On-device TTS with native SDKs. 28 voices in 10 languages, voice cloning from about 10 s of audio, and a local server with an OpenAI-compatible speech endpoint. PyTorch, ONNX Runtime and CoreML backends; the Swift, Go, Rust and TypeScript ports run without Python. Apache-2.0, derived from Chatterbox. | Python / Swift / Go / Rust / TypeScript |
Dedicated resources and examples focusing on the latest in voice replication and advanced synthetic voice generation.
Hugging Face has emerged as a central ecosystem for sharing, discovering, and experimenting with a vast array of pretrained Text-to-Speech models. Explore their extensive collection for diverse applications and research.
Stay updated with the latest breakthroughs and discussions in the TTS community.
A collection of influential code repositories and product demonstrations showcasing various Text-to-Speech implementations and their output quality. Sorted by Year of Launch (Descending).
| Project/Samples | Pretrained Models | Code Link | Paper/Arxiv ID | Output Quality | Year of Launch | Description |
|---|---|---|---|---|---|---|
| Gradbot Demos | -- | Code | Codebase | A | 2026 | Eight voice agent demos (banking, hotel booking, 3D game NPCs) built on Gradium's real-time TTS/STT APIs. |
| Fish Speech v1.5 | -- | Code | Codebase | A+ | 2026 | SOTA multilingual, multi-speaker model with superior naturalness. |
| GLM-4-Voice Samples | -- | Code | Codebase | A | 2025 | End-to-end voice model with expressive bilingual dialogue. |
| Kokoro-82M Samples | -- | Code | -- | A | 2025 | Ultra-efficient CPU-based model with studio-quality output. |
| ChatTTS Samples | -- | Code | 2406.03807 | A+ | 2024 | Conversational dialogue speech synthesis with prosodic laughs and pauses. |
| CosyVoice Samples | -- | Code | 2407.05407 | A+ | 2024 | Multilingual voice generation with emotional control and multi-dialect support. |
| F5-TTS Samples | -- | Code | 2410.06885 | A | 2024 | Diffusion-based zero-shot cloning with impressive prosody. |
| GPT-SoVITS Samples | -- | Code | Codebase | A+ | 2024 | Powerful few-shot voice cloning requiring only 5 seconds of sample audio. |
| MaskGCT Samples | -- | Code | 2409.00750 | A | 2024 | Non-autoregressive zero-shot TTS using masked generative codec transformers. |
| MeloTTS Samples | -- | Code | Codebase | B | 2024 | Multilingual, multi-speaker TTS model for high-quality audio generation. |
| Parler-TTS Samples | -- | Code | 2402.01912 | B | 2024 | Samples from a lightweight model producing natural-sounding speech. |
| VoiceCraft Samples | -- | Code | 2403.16973 | A | 2024 | Zero-shot speech editing and neural synthesis with token infilling. |
| Bark Samples (Suno.ai) | -- | Code | -- | A | 2023 | Samples from Suno's expressive text-to-audio model, including non-speech sounds. |
| StyleTTS 2 Samples | -- | Code | 2306.07691 | A | 2023 | Human-level TTS with style diffusion and adversarial training. |
| VALL-E X Samples | -- | Code | 2303.03926 | A | 2023 | Cross-lingual zero-shot speech synthesis and voice cloning. |
| XTTS-v2 Samples | -- | Code | 2309.02055 | A | 2023 | Demonstrations of Coqui's advanced voice cloning with emotion transfer. |
| rayhane's Tacotron2 Samples | -- | -- | -- | D | 2019 | Audio samples from an early Tacotron 2 implementation. |
| Google Tacotron + Style Transfer Sample (Official) | -- | -- | 1803.09047 | A | 2018 | Official samples showcasing prosody and style transfer with Tacotron. |
| Kyubyong's DC-TTS on Nick Dataset Samples | -- | -- | -- | D | 2018 | DC-TTS samples generated from the Nick dataset. |
| Kyubyong's Expressive Tacotron Samples | -- | Code | 1803.09047 | D | 2018 | Samples demonstrating expressive speech synthesis with Tacotron. |
| Kyubyong's Tacotron on LJ Dataset Samples | Download model | -- | -- | D | 2018 | Audio generated from Tacotron trained on the LJSpeech dataset. |
| Kyubyong's Tacotron on Nick Dataset Samples | -- | -- | -- | D | 2018 | Tacotron samples from the Nick dataset. |
| Kyubyong's Tacotron on Web Dataset Samples | Download model | -- | -- | D | 2018 | Tacotron speech output from the Web dataset. |
| mazzzystar's Tacotron-WaveRNN Samples | Get Model | Code | -- | A | 2018 | Demonstrations from a Tacotron and WaveRNN hybrid model. |
| NVIDIA's Tacotron2 + WaveGlow Samples | Download Model | Code | -- | A | 2018 | Combined high-quality speech synthesis from Tacotron 2 and WaveGlow. |
| NVIDIA's WaveGlow Samples | Download Model | Code | 1811.00002 | A | 2018 | High-fidelity audio generated by NVIDIA's WaveGlow vocoder. |
| syang1993's Tacotron + Style Transfer Samples | Model ErnstTmp (232k iter) | -- | 1803.09047 and 1803.09017 | C | 2018 | Samples demonstrating Tacotron with global style tokens for voice style transfer. |
| andabi's Deep Voice Conversion | -- | -- | -- | D | 2017 | Demonstrations of deep voice conversion techniques. |
| Baidu's Deep Voice Samples (Official) | -- | -- | -- | D | 2017 | Official audio demonstrations from Baidu's Deep Voice project. |
| Baidu's Deep Voice 3 Samples (Official) | -- | -- | 1710.07654 | B | 2017 | Official samples from Deep Voice 3, showcasing advanced speech synthesis. |
| DeepMind Neural Discrete Representation Learning Samples (Official) | -- | -- | 1711.00937 | B | 2017 | Samples demonstrating speech generated using VQ-VAE for neural discrete representation learning. |
| dhgrs's Implementation of Neural Discrete Representation Learning Samples | Download Model | Code | 1711.00937 | D | 2017 | Audio generated using a Chainer implementation of VQ-VAE for speech. |
| Facebook Loop Samples (Official) | Get model | -- | -- | D | 2017 | Official audio samples from Facebook's Loop project. |
| Google Tacotron2 Samples (Official) | -- | -- | 1712.05884 | A | 2017 | Official, high-quality audio samples from the groundbreaking Tacotron 2 model. |
| keithito's Tacotron Samples | Get model | -- | -- | D | 2017 | Audio samples from keithito's Tacotron implementation. |
| Kyubyong's DC-TTS Kate Samples | -- | -- | -- | D | 2017 | DC-TTS samples featuring the "Kate" voice. |
| Kyubyong's DC-TTS on LJ Dataset Samples | Get model | -- | -- | D | 2017 | DC-TTS generated speech from the LJSpeech dataset. |
| mazzzystar's RandomCNN Voice Transfer | -- | -- | 1712.08363 | D | 2017 | Speech conversion samples using Random CNNs. |
| r9y9's Wavenet Vocoder Tacotron2 Samples | Download Tacotron2 model - Download Wavenet model - Get models | -- | 1712.05884 and 1611.09482 | B | 2017 | Samples from a Tacotron 2 and WaveNet vocoder combination. |
| Griffin-Lim Samples | -- | -- | -- | A | 1984 | Classic samples from the Griffin-Lim algorithm for spectrogram inversion. |
Ongoing projects and cutting-edge research shaping the next generation of AI voice synthesis.
If I missed your output sample/demo in this consolidation, just add and send a pull request. I will be more than happy to add it. Thanks!
Practical guides and interactive notebooks for experimenting with Text-to-Speech models.
Visual demonstrations of advanced Text-to-Speech and voice cloning in action.
Broader projects and research efforts that contribute to the Text-to-Speech ecosystem.
Explore influential academic papers and preprints in the field of Text-to-Speech and voice AI.
Connect with the community, get support, and stay informed about the latest in TTS.
Explore how Text-to-Speech and Voice Cloning are being used across industries:
As of 2026, Kokoro-82M is widely considered the best for CPU-based local inference due to its studio quality and small footprint. For high-fidelity and expressive speech, F5-TTS and Fish Speech are leading the way in naturalness.
Yes! Projects like Coqui XTTS-v2, GPT-SoVITS, and OpenVoice offer high-quality voice cloning for free. If you are looking for local-first alternatives, check out F5-TTS and CosyVoice.
To achieve sub-200ms latency, it is recommended to use Deepgram Aura, Cartesia Sonic, Smallest.ai, or optimized local models like Piper (C++ implementation) and Kokoro-82M with ONNX runtime.
Many models listed here (like OpenAI TTS, ElevenLabs, and Azure AI Speech) have clear commercial tiers. For open-source models, look for those with MIT or Apache 2.0 licenses, such as Piper and Kokoro.
If you find this collection of Text-to-Speech resources helpful, or if it has saved you time and effort in your AI voice generation endeavors, please consider sponsoring the development. Your support helps maintain the project, add new cutting-edge models and tools, and keep this initiative open-source and accessible to everyone.
Sponsor @ishandutta2007 on GitHub
Every contribution, no matter how small, makes a huge difference in advancing the Text-to-Speech landscape! 🙏
This project is licensed under the MIT License - see the LICENSE file for details.
🎤 A curated list of the latest and most influential tools, models, and resources in the Text-to-Speech sector. 🌟 Star if you like it! 🌟
194
119 commits
updated Sep 14, 2026
Welcome to the most comprehensive, meticulously curated, and continuously updated list of Text-to-Speech (TTS) resources. Whether you are looking for the best open-source TTS models of 2026, searching for low-latency TTS APIs for AI agents, or exploring high-fidelity voice cloning for content creation, you've found the right place.
[!TIP] Looking for the best ElevenLabs alternatives? This repository tracks the rapidly evolving landscape of both commercial SaaS and local-first neural speech synthesis.
Text-to-Speech technology has moved beyond robotic voices. Today, it powers:
The landscape of AI voice synthesis has shifted from basic concatenation to advanced Generative Speech Models. Key highlights:
Leading platforms offering robust, scalable, and high-quality Text-to-Speech APIs and services for various applications. Sorted by Company Scale / Valuation (Descending).
| 🛠️ Service/Model | 🏢 Organization | 🌟 Key Features | 💰 Min. Monthly Subscription | 👥 Company Size | 🔗 Link |
|---|---|---|---|---|---|
| NVIDIA NeMo | NVIDIA | Platform for building, training, and deploying generative AI models, including TTS and ASR. | Free (API Credits) / Enterprise | $3.3T+ (Market Cap) | NVIDIA NeMo |
| Azure AI Speech | Microsoft | High-quality neural voices with advanced fine-tuning, emotion, and enterprise scalability. | Free (0.5M chars/mo) / PAYG | $3.2T+ (Market Cap) | Azure AI Speech |
| Google Cloud TTS | Powerful TTS API with a large variety of natural-sounding voices and extensive customization. | Free (1M+ chars/mo) / PAYG | $2.2T+ (Market Cap) | Google Cloud TTS | |
| AWS Polly | Amazon | Generative, Neural and Standard TTS voices with deep AWS ecosystem integration. | Free (1M+ chars/mo) / PAYG | $1.9T+ (Market Cap) | AWS Polly |
| OpenAI TTS | OpenAI | High-quality, real-time streaming TTS models for applications requiring natural AI voices. | Pay-as-you-go ($5 free credit) | $852B (Valuation) | OpenAI TTS |
| ElevenLabs | ElevenLabs | State-of-the-art AI voice generator offering realistic voices, voice cloning, and AI dubbing. | Free (10k chars/mo) / $5 | $11B (Valuation) | ElevenLabs |
| Speechify | Speechify | Highly popular consumer text-to-speech and developer Voice API with natural and premium voices. | Free (Basic Voices) / $139/yr | $1.5B (Valuation) | Speechify |
| Deepgram Aura | Deepgram | Specializing in low-latency TTS designed for real-time conversational AI and virtual interactions. | Free ($200 credit) / PAYG | $1.2B (Valuation) | Deepgram Aura |
| Inworld AI | Inworld AI | Character-driven conversational voice engine and real-time speech generation for interactive agents. | Free (Basic / 100k API credits) / $20 | $500M (Valuation) | Inworld AI |
| Hume AI (EVI) | Hume AI | Empathic Voice Interface with emotional prosody detection and low-latency expressive speech. | Free ($10 credit) / PAYG ($0.036/min) | $250M (Valuation) | Hume AI |
| Cartesia Sonic | Cartesia | Sub-100ms ultra-low latency TTS designed for real-time AI agents. | Free (API Credits) / PAYG | $200M (Valuation) | Cartesia |
| Gradium | Gradium | Real-time TTS and STT for voice agents. 158ms P50 time-to-first-audio, streaming and instant voice cloning. | Free (45k credits) / $13 | $100M (Seed Raised) | Gradium |
| Murf.ai | Murf.ai | AI voiceovers with a built-in video editor, ideal for creators and presentations. | Free (10 mins total) / $19 | $46M (Valuation) | Murf.ai |
| WellSaid Labs | WellSaid Labs | Enterprise AI voice platform with natural studio voices, fine phonetic control, and brand voice avatars. | Free (7-day trial / 50 clips) / $49 | $40M (Valuation) | WellSaid Labs |
| Resemble AI | Resemble AI | High-fidelity voice cloning, neural watermarking, deepfake detection, and real-time speech synthesis. | Free (Free trial) / $29 | $30M (Valuation) | Resemble AI |
| LMNT | LMNT | Lightning-fast TTS API with exceptional naturalness, great for interactive voice apps. | Free (API Credits) / PAYG | $21M (Estimated) | LMNT |
| Lovo.ai (Genny) | LOVO | Next-gen AI voice generator with 500+ voices in 100+ languages and granular pitch/emphasis controls. | Free (14-day Pro trial) / $29 | $20M (Estimated) | Lovo.ai |
| Play.ht | Play.ht | Professional AI voices and "Ultra-Realistic" studio editor for long-form content. | Free (12.5k chars) / $39 | $15M (Valuation) | Play.ht |
| Smallest.ai (Waves) | Smallest.ai | Lightning-fast TTS API with sub-100ms latency designed for real-time conversational agents and voice bots. | Free (10k chars/mo) / PAYG ($0.08/1k chars) | $12M (Valuation) | Smallest.ai |
| Rime Labs | Rime | Ultra-fast expressive voice API built specifically for interactive conversational AI applications. | Free ($5 credit) / PAYG | $10M (Estimated) | Rime |
| Soniox TTS | Soniox | Real-time streaming TTS API for conversational AI voice agents in 60+ languages with multilingual voices. | Pay-as-you-go (~$0.70/hr) | $10M (Estimated) | Soniox |
| Neets.ai | Neets | Extremely fast and affordable TTS APIs starting at $0.0004 per 1k characters. | Free (API Credits) / PAYG | $5M (Estimated) | Neets.ai |
| Gandr | Gandr | TTS API for voice agents. Word error rate 1.98 percent against a 2.17 percent human reference on the same scorer, one voice in 23 languages, every render watermarked. Python/JS SDKs, LiveKit plugin, MCP server. | From $10/mo. Flat-rate unmetered streams from $150/mo | <$1M (Indie) | Gandr |
| Spokio | Spokio | Offline macOS text-to-speech app with local voice cloning, batch export, and no cloud uploads. | Free (API Credits) / Enterprise | <$1M (Indie) | Spokio |
| PHANTOM VOICES | PHANTOM VOICES | 10 free professional AI voice clones via public REST API. Zero cost, commercial rights cleared. 29 platform configs (Vapi, Retell AI, etc). Multilingual (9+ languages). AI-powered recommendation. | Free (API Credits) / Enterprise | <$1M (Indie) | PHANTOM VOICES |
| RunAPI ElevenLabs SDK | RunAPI | Multi-language SDKs for ElevenLabs text-to-speech, dialogue generation, sound effects, transcription, and audio isolation workflows. | Pay-as-you-go | <$1M (Indie) | RunAPI ElevenLabs SDK |
| Audexum | Audexum | Text-to-speech and speech-to-text in one API: 43 voices, 32 TTS languages, 25 STT languages. ElevenLabs-compatible endpoint, so switching is a base-URL change. EU-hosted, every output watermarked. | Free (30k credits at signup, 3k/mo after) / EUR 4 | <$1M (Indie) | Audexum |
If you are looking for free text-to-speech models for commercial use or want to run TTS locally on a CPU, these open-source projects provide the best balance of quality and privacy. Sorted by GitHub Star Counts (Descending).
| 🛠️ Service/Model | 🏢 Organization | 🌟 Key Features | 🗣️ Primary Language | 📁 Github_Repository |
|---|---|---|---|---|
| 🐸 Coqui TTS | Coqui (community) | Supports 1100+ languages, zero-shot voice cloning, and fine-tuning. Note: Coqui AI (the company) shut down in late 2023; actively maintained by the community at idiap/coqui-ai-TTS. | Python / Multilingual | |
| GPT-SoVITS | RVC-Boss | Zero-shot & few-shot voice cloning. Requires only 5 seconds of sample audio for cross-lingual synthesis. | Python / PyTorch | |
| Bark | Suno | Transformer-based text-to-audio model capable of highly expressive speech, music, laughs, and sighs. | Python / PyTorch | |
| ChatTTS | 2noise | Conversational text-to-speech model specially optimized for dialogue and natural conversational flow. | Python | |
| OpenVoice | MyShell | Highly versatile and instant voice cloning that requires only a short audio clip. | Python | |
| Fish Speech | Fish Audio | SOTA multilingual, multi-speaker model with superior naturalness. | Python | |
| Chatterbox | Resemble AI | Advanced neural voice synthesis with emotion control and high-fidelity cloning. | Python | |
| CosyVoice | Alibaba | Excellent multilingual and zero-shot voice cloning model capable of high fidelity. | Python | |
| KittenTTS | KittenML | ONNX-based library for low-latency TTS without requiring a GPU. | Python / ONNX | |
| F5-TTS | SWivid | Flow Matching TTS. Incredible naturalness and prosody using DiT architectures. | Python | |
| Tortoise-TTS | James Betker | Powerful multi-voice TTS system known for its exceptional voice cloning capabilities. | Python | |
| VALL-E-X | Plachtaa | Open-source implementation of Microsoft's VALL-E X for zero-shot cross-lingual voice cloning. | Python / PyTorch | |
| Piper | Rhasspy | Fastest local TTS. Optimized for low-end hardware and offline use. | C++ / Python | |
| Amphion | Amphion | Open-source audio, music and speech generation toolkit containing multiple SOTA TTS models. | Python | |
| VoiceCraft | Jason Li et al. | Token-infilling neural codec model for zero-shot speech editing and TTS synthesis. | Python / PyTorch | |
| Kokoro-82M | Hexgrad | Best SOTA CPU TTS. Ultra-fast, studio quality, only 82M parameters. | ONNX / Python | |
| StyleTTS 2 | yl4579 | Human-level TTS. Uses style diffusion, adversarial training, and large SLMs without phoneme duration models. | Python / PyTorch | |
| MeloTTS | MyShell | Ultra-fast multilingual TTS running smoothly on CPU across English, Spanish, French, Chinese, Japanese, Korean. | Python / PyTorch | |
| sherpa-onnx | Next-gen Kaldi | Offline multi-platform speech synthesis engine supporting VITS, Piper, and Kokoro on embedded/mobile/desktop. | C++ / Python / Go | |
| Parler-TTS | Hugging Face | Lightweight, controllable speech generation with high naturalness. | Python | HF |
| GLM-4-Voice | Zhipu AI | End-to-end voice model supporting real-time speech generation, emotion alteration, and bilingual dialogue. | Python / PyTorch | |
| Matcha-TTS | Shivam Mehta | Fast TTS architecture employing conditional flow matching, producing highly natural output. | Python | |
| LocalMode | LocalMode | In-browser TTS. Runs Kokoro (29 voices) and other AI models 100% in the browser via WebGPU/WASM. No server, no API keys, offline after first load. | JavaScript / TypeScript | |
| Vocello | PowerBeef | Native Mac & iPhone app. Qwen3-TTS with preset speakers, natural-language voice design, and voice cloning. Runs entirely on Apple Silicon with no Python runtime, faster than realtime on an 8 GB M2. | Swift / MLX | |
| loudkit | LoudReader | On-device TTS with native SDKs. 28 voices in 10 languages, voice cloning from about 10 s of audio, and a local server with an OpenAI-compatible speech endpoint. PyTorch, ONNX Runtime and CoreML backends; the Swift, Go, Rust and TypeScript ports run without Python. Apache-2.0, derived from Chatterbox. | Python / Swift / Go / Rust / TypeScript |
Dedicated resources and examples focusing on the latest in voice replication and advanced synthetic voice generation.
Hugging Face has emerged as a central ecosystem for sharing, discovering, and experimenting with a vast array of pretrained Text-to-Speech models. Explore their extensive collection for diverse applications and research.
Stay updated with the latest breakthroughs and discussions in the TTS community.
A collection of influential code repositories and product demonstrations showcasing various Text-to-Speech implementations and their output quality. Sorted by Year of Launch (Descending).
| Project/Samples | Pretrained Models | Code Link | Paper/Arxiv ID | Output Quality | Year of Launch | Description |
|---|---|---|---|---|---|---|
| Gradbot Demos | -- | Code | Codebase | A | 2026 | Eight voice agent demos (banking, hotel booking, 3D game NPCs) built on Gradium's real-time TTS/STT APIs. |
| Fish Speech v1.5 | -- | Code | Codebase | A+ | 2026 | SOTA multilingual, multi-speaker model with superior naturalness. |
| GLM-4-Voice Samples | -- | Code | Codebase | A | 2025 | End-to-end voice model with expressive bilingual dialogue. |
| Kokoro-82M Samples | -- | Code | -- | A | 2025 | Ultra-efficient CPU-based model with studio-quality output. |
| ChatTTS Samples | -- | Code | 2406.03807 | A+ | 2024 | Conversational dialogue speech synthesis with prosodic laughs and pauses. |
| CosyVoice Samples | -- | Code | 2407.05407 | A+ | 2024 | Multilingual voice generation with emotional control and multi-dialect support. |
| F5-TTS Samples | -- | Code | 2410.06885 | A | 2024 | Diffusion-based zero-shot cloning with impressive prosody. |
| GPT-SoVITS Samples | -- | Code | Codebase | A+ | 2024 | Powerful few-shot voice cloning requiring only 5 seconds of sample audio. |
| MaskGCT Samples | -- | Code | 2409.00750 | A | 2024 | Non-autoregressive zero-shot TTS using masked generative codec transformers. |
| MeloTTS Samples | -- | Code | Codebase | B | 2024 | Multilingual, multi-speaker TTS model for high-quality audio generation. |
| Parler-TTS Samples | -- | Code | 2402.01912 | B | 2024 | Samples from a lightweight model producing natural-sounding speech. |
| VoiceCraft Samples | -- | Code | 2403.16973 | A | 2024 | Zero-shot speech editing and neural synthesis with token infilling. |
| Bark Samples (Suno.ai) | -- | Code | -- | A | 2023 | Samples from Suno's expressive text-to-audio model, including non-speech sounds. |
| StyleTTS 2 Samples | -- | Code | 2306.07691 | A | 2023 | Human-level TTS with style diffusion and adversarial training. |
| VALL-E X Samples | -- | Code | 2303.03926 | A | 2023 | Cross-lingual zero-shot speech synthesis and voice cloning. |
| XTTS-v2 Samples | -- | Code | 2309.02055 | A | 2023 | Demonstrations of Coqui's advanced voice cloning with emotion transfer. |
| rayhane's Tacotron2 Samples | -- | -- | -- | D | 2019 | Audio samples from an early Tacotron 2 implementation. |
| Google Tacotron + Style Transfer Sample (Official) | -- | -- | 1803.09047 | A | 2018 | Official samples showcasing prosody and style transfer with Tacotron. |
| Kyubyong's DC-TTS on Nick Dataset Samples | -- | -- | -- | D | 2018 | DC-TTS samples generated from the Nick dataset. |
| Kyubyong's Expressive Tacotron Samples | -- | Code | 1803.09047 | D | 2018 | Samples demonstrating expressive speech synthesis with Tacotron. |
| Kyubyong's Tacotron on LJ Dataset Samples | Download model | -- | -- | D | 2018 | Audio generated from Tacotron trained on the LJSpeech dataset. |
| Kyubyong's Tacotron on Nick Dataset Samples | -- | -- | -- | D | 2018 | Tacotron samples from the Nick dataset. |
| Kyubyong's Tacotron on Web Dataset Samples | Download model | -- | -- | D | 2018 | Tacotron speech output from the Web dataset. |
| mazzzystar's Tacotron-WaveRNN Samples | Get Model | Code | -- | A | 2018 | Demonstrations from a Tacotron and WaveRNN hybrid model. |
| NVIDIA's Tacotron2 + WaveGlow Samples | Download Model | Code | -- | A | 2018 | Combined high-quality speech synthesis from Tacotron 2 and WaveGlow. |
| NVIDIA's WaveGlow Samples | Download Model | Code | 1811.00002 | A | 2018 | High-fidelity audio generated by NVIDIA's WaveGlow vocoder. |
| syang1993's Tacotron + Style Transfer Samples | Model ErnstTmp (232k iter) | -- | 1803.09047 and 1803.09017 | C | 2018 | Samples demonstrating Tacotron with global style tokens for voice style transfer. |
| andabi's Deep Voice Conversion | -- | -- | -- | D | 2017 | Demonstrations of deep voice conversion techniques. |
| Baidu's Deep Voice Samples (Official) | -- | -- | -- | D | 2017 | Official audio demonstrations from Baidu's Deep Voice project. |
| Baidu's Deep Voice 3 Samples (Official) | -- | -- | 1710.07654 | B | 2017 | Official samples from Deep Voice 3, showcasing advanced speech synthesis. |
| DeepMind Neural Discrete Representation Learning Samples (Official) | -- | -- | 1711.00937 | B | 2017 | Samples demonstrating speech generated using VQ-VAE for neural discrete representation learning. |
| dhgrs's Implementation of Neural Discrete Representation Learning Samples | Download Model | Code | 1711.00937 | D | 2017 | Audio generated using a Chainer implementation of VQ-VAE for speech. |
| Facebook Loop Samples (Official) | Get model | -- | -- | D | 2017 | Official audio samples from Facebook's Loop project. |
| Google Tacotron2 Samples (Official) | -- | -- | 1712.05884 | A | 2017 | Official, high-quality audio samples from the groundbreaking Tacotron 2 model. |
| keithito's Tacotron Samples | Get model | -- | -- | D | 2017 | Audio samples from keithito's Tacotron implementation. |
| Kyubyong's DC-TTS Kate Samples | -- | -- | -- | D | 2017 | DC-TTS samples featuring the "Kate" voice. |
| Kyubyong's DC-TTS on LJ Dataset Samples | Get model | -- | -- | D | 2017 | DC-TTS generated speech from the LJSpeech dataset. |
| mazzzystar's RandomCNN Voice Transfer | -- | -- | 1712.08363 | D | 2017 | Speech conversion samples using Random CNNs. |
| r9y9's Wavenet Vocoder Tacotron2 Samples | Download Tacotron2 model - Download Wavenet model - Get models | -- | 1712.05884 and 1611.09482 | B | 2017 | Samples from a Tacotron 2 and WaveNet vocoder combination. |
| Griffin-Lim Samples | -- | -- | -- | A | 1984 | Classic samples from the Griffin-Lim algorithm for spectrogram inversion. |
Ongoing projects and cutting-edge research shaping the next generation of AI voice synthesis.
If I missed your output sample/demo in this consolidation, just add and send a pull request. I will be more than happy to add it. Thanks!
Practical guides and interactive notebooks for experimenting with Text-to-Speech models.
Visual demonstrations of advanced Text-to-Speech and voice cloning in action.
Broader projects and research efforts that contribute to the Text-to-Speech ecosystem.
Explore influential academic papers and preprints in the field of Text-to-Speech and voice AI.
Connect with the community, get support, and stay informed about the latest in TTS.
Explore how Text-to-Speech and Voice Cloning are being used across industries:
As of 2026, Kokoro-82M is widely considered the best for CPU-based local inference due to its studio quality and small footprint. For high-fidelity and expressive speech, F5-TTS and Fish Speech are leading the way in naturalness.
Yes! Projects like Coqui XTTS-v2, GPT-SoVITS, and OpenVoice offer high-quality voice cloning for free. If you are looking for local-first alternatives, check out F5-TTS and CosyVoice.
To achieve sub-200ms latency, it is recommended to use Deepgram Aura, Cartesia Sonic, Smallest.ai, or optimized local models like Piper (C++ implementation) and Kokoro-82M with ONNX runtime.
Many models listed here (like OpenAI TTS, ElevenLabs, and Azure AI Speech) have clear commercial tiers. For open-source models, look for those with MIT or Apache 2.0 licenses, such as Piper and Kokoro.
If you find this collection of Text-to-Speech resources helpful, or if it has saved you time and effort in your AI voice generation endeavors, please consider sponsoring the development. Your support helps maintain the project, add new cutting-edge models and tools, and keep this initiative open-source and accessible to everyone.
Sponsor @ishandutta2007 on GitHub
Every contribution, no matter how small, makes a huge difference in advancing the Text-to-Speech landscape! 🙏
This project is licensed under the MIT License - see the LICENSE file for details.