ishandutta2007/Awesome-Text-to-Speech

🎤 A curated list of the latest and most influential tools, models, and resources in the Text-to-Speech sector. 🌟 Star if you like it! 🌟

194

119 commits

updated Sep 14, 2026

See the code

README

Awesome Text-to-Speech Banner

Awesome Text-to-Speech (TTS) 🗣️: Best AI Voice Generation Models & Tools 2026

AwesomeDiscord Awesome GitHub stars GitHub forks GitHub license Python Version PRs Welcome GitHub Sponsors Last Commit Contributors Follow on Twitter GitHub followers


🚀 The Ultimate Guide to AI Voice Generation, Speech Synthesis, and Neural Voice Cloning

Welcome to the most comprehensive, meticulously curated, and continuously updated list of Text-to-Speech (TTS) resources. Whether you are looking for the best open-source TTS models of 2026, searching for low-latency TTS APIs for AI agents, or exploring high-fidelity voice cloning for content creation, you've found the right place.

[!TIP] Looking for the best ElevenLabs alternatives? This repository tracks the rapidly evolving landscape of both commercial SaaS and local-first neural speech synthesis.


🗺️ Quick Navigation


🔊 Why Explore Text-to-Speech?

Text-to-Speech technology has moved beyond robotic voices. Today, it powers:

  • Accessibility First: High-quality screen readers for the visually impaired.
  • Automated Content Creation: Realistic voiceovers for YouTube, podcasts, and e-learning.
  • Next-Gen AI Agents: Real-time conversational AI with human-like prosody.
  • Multilingual Support: Instant translation and dubbing for global reach.
  • Personalization: Custom voice clones for gaming and virtual assistants.

📈 Current State of Text-to-Speech (2026 Update)

The landscape of AI voice synthesis has shifted from basic concatenation to advanced Generative Speech Models. Key highlights:

  • Hyper-realistic and Natural Speech Synthesis: Innovations in deep learning and neural network architectures have led to highly natural, expressive, and emotionally nuanced synthetic voices. 🎤
  • Next-Generation Architectures: The adoption of State Space Models (SSMs), Diffusion Models, and advanced transformer-based architectures is offering superior performance, efficiency, and voice quality in speech generation. 🧠
  • Real-time Conversational AI: Significant advancements in reducing latency now enable real-time TTS, making conversational AI, virtual assistants, and live dubbing more natural and responsive. ⚡
  • Advanced Voice Cloning and Style Transfer: Cutting-edge techniques allow for high-fidelity voice cloning from minimal audio samples and the transfer of speaking style and emotion across different voices. 🎭
  • Multilingual and Cross-Lingual TTS: Models are increasingly capable of generating speech in numerous languages with accurate pronunciation and intonation, breaking down language barriers.

📚 Comprehensive List of Text-to-Speech (TTS) Resources 🌐

☁️ Cloud-based & Commercial AI Voice Generation Platforms

Leading platforms offering robust, scalable, and high-quality Text-to-Speech APIs and services for various applications. Sorted by Company Scale / Valuation (Descending).

🛠️ Service/Model🏢 Organization🌟 Key Features💰 Min. Monthly Subscription👥 Company Size🔗 Link
NVIDIA NeMoNVIDIAPlatform for building, training, and deploying generative AI models, including TTS and ASR.Free (API Credits) / Enterprise$3.3T+ (Market Cap)NVIDIA NeMo
Azure AI SpeechMicrosoftHigh-quality neural voices with advanced fine-tuning, emotion, and enterprise scalability.Free (0.5M chars/mo) / PAYG$3.2T+ (Market Cap)Azure AI Speech
Google Cloud TTSGooglePowerful TTS API with a large variety of natural-sounding voices and extensive customization.Free (1M+ chars/mo) / PAYG$2.2T+ (Market Cap)Google Cloud TTS
AWS PollyAmazonGenerative, Neural and Standard TTS voices with deep AWS ecosystem integration.Free (1M+ chars/mo) / PAYG$1.9T+ (Market Cap)AWS Polly
OpenAI TTSOpenAIHigh-quality, real-time streaming TTS models for applications requiring natural AI voices.Pay-as-you-go ($5 free credit)$852B (Valuation)OpenAI TTS
ElevenLabsElevenLabsState-of-the-art AI voice generator offering realistic voices, voice cloning, and AI dubbing.Free (10k chars/mo) / $5$11B (Valuation)ElevenLabs
SpeechifySpeechifyHighly popular consumer text-to-speech and developer Voice API with natural and premium voices.Free (Basic Voices) / $139/yr$1.5B (Valuation)Speechify
Deepgram AuraDeepgramSpecializing in low-latency TTS designed for real-time conversational AI and virtual interactions.Free ($200 credit) / PAYG$1.2B (Valuation)Deepgram Aura
Inworld AIInworld AICharacter-driven conversational voice engine and real-time speech generation for interactive agents.Free (Basic / 100k API credits) / $20$500M (Valuation)Inworld AI
Hume AI (EVI)Hume AIEmpathic Voice Interface with emotional prosody detection and low-latency expressive speech.Free ($10 credit) / PAYG ($0.036/min)$250M (Valuation)Hume AI
Cartesia SonicCartesiaSub-100ms ultra-low latency TTS designed for real-time AI agents.Free (API Credits) / PAYG$200M (Valuation)Cartesia
GradiumGradiumReal-time TTS and STT for voice agents. 158ms P50 time-to-first-audio, streaming and instant voice cloning.Free (45k credits) / $13$100M (Seed Raised)Gradium
Murf.aiMurf.aiAI voiceovers with a built-in video editor, ideal for creators and presentations.Free (10 mins total) / $19$46M (Valuation)Murf.ai
WellSaid LabsWellSaid LabsEnterprise AI voice platform with natural studio voices, fine phonetic control, and brand voice avatars.Free (7-day trial / 50 clips) / $49$40M (Valuation)WellSaid Labs
Resemble AIResemble AIHigh-fidelity voice cloning, neural watermarking, deepfake detection, and real-time speech synthesis.Free (Free trial) / $29$30M (Valuation)Resemble AI
LMNTLMNTLightning-fast TTS API with exceptional naturalness, great for interactive voice apps.Free (API Credits) / PAYG$21M (Estimated)LMNT
Lovo.ai (Genny)LOVONext-gen AI voice generator with 500+ voices in 100+ languages and granular pitch/emphasis controls.Free (14-day Pro trial) / $29$20M (Estimated)Lovo.ai
Play.htPlay.htProfessional AI voices and "Ultra-Realistic" studio editor for long-form content.Free (12.5k chars) / $39$15M (Valuation)Play.ht
Smallest.ai (Waves)Smallest.aiLightning-fast TTS API with sub-100ms latency designed for real-time conversational agents and voice bots.Free (10k chars/mo) / PAYG ($0.08/1k chars)$12M (Valuation)Smallest.ai
Rime LabsRimeUltra-fast expressive voice API built specifically for interactive conversational AI applications.Free ($5 credit) / PAYG$10M (Estimated)Rime
Soniox TTSSonioxReal-time streaming TTS API for conversational AI voice agents in 60+ languages with multilingual voices.Pay-as-you-go (~$0.70/hr)$10M (Estimated)Soniox
Neets.aiNeetsExtremely fast and affordable TTS APIs starting at $0.0004 per 1k characters.Free (API Credits) / PAYG$5M (Estimated)Neets.ai
GandrGandrTTS API for voice agents. Word error rate 1.98 percent against a 2.17 percent human reference on the same scorer, one voice in 23 languages, every render watermarked. Python/JS SDKs, LiveKit plugin, MCP server.From $10/mo. Flat-rate unmetered streams from $150/mo<$1M (Indie)Gandr
SpokioSpokioOffline macOS text-to-speech app with local voice cloning, batch export, and no cloud uploads.Free (API Credits) / Enterprise<$1M (Indie)Spokio
PHANTOM VOICESPHANTOM VOICES10 free professional AI voice clones via public REST API. Zero cost, commercial rights cleared. 29 platform configs (Vapi, Retell AI, etc). Multilingual (9+ languages). AI-powered recommendation.Free (API Credits) / Enterprise<$1M (Indie)PHANTOM VOICES
RunAPI ElevenLabs SDKRunAPIMulti-language SDKs for ElevenLabs text-to-speech, dialogue generation, sound effects, transcription, and audio isolation workflows.Pay-as-you-go<$1M (Indie)RunAPI ElevenLabs SDK
AudexumAudexumText-to-speech and speech-to-text in one API: 43 voices, 32 TTS languages, 25 STT languages. ElevenLabs-compatible endpoint, so switching is a base-URL change. EU-hosted, every output watermarked.Free (30k credits at signup, 3k/mo after) / EUR 4<$1M (Indie)Audexum

🏗️ Open-Source Text-to-Speech Libraries & Local-First Projects

If you are looking for free text-to-speech models for commercial use or want to run TTS locally on a CPU, these open-source projects provide the best balance of quality and privacy. Sorted by GitHub Star Counts (Descending).

Sound Wave Animation
🛠️ Service/Model🏢 Organization🌟 Key Features🗣️ Primary Language📁 Github_Repository
🐸 Coqui TTSCoqui (community)Supports 1100+ languages, zero-shot voice cloning, and fine-tuning. Note: Coqui AI (the company) shut down in late 2023; actively maintained by the community at idiap/coqui-ai-TTS.Python / MultilingualGitHub stars
GPT-SoVITSRVC-BossZero-shot & few-shot voice cloning. Requires only 5 seconds of sample audio for cross-lingual synthesis.Python / PyTorchGitHub stars
BarkSunoTransformer-based text-to-audio model capable of highly expressive speech, music, laughs, and sighs.Python / PyTorchGitHub stars
ChatTTS2noiseConversational text-to-speech model specially optimized for dialogue and natural conversational flow.PythonGitHub stars
OpenVoiceMyShellHighly versatile and instant voice cloning that requires only a short audio clip.PythonGitHub stars
Fish SpeechFish AudioSOTA multilingual, multi-speaker model with superior naturalness.PythonGitHub stars
ChatterboxResemble AIAdvanced neural voice synthesis with emotion control and high-fidelity cloning.PythonGitHub stars
CosyVoiceAlibabaExcellent multilingual and zero-shot voice cloning model capable of high fidelity.PythonGitHub stars
KittenTTSKittenMLONNX-based library for low-latency TTS without requiring a GPU.Python / ONNXGitHub stars
F5-TTSSWividFlow Matching TTS. Incredible naturalness and prosody using DiT architectures.PythonGitHub stars
Tortoise-TTSJames BetkerPowerful multi-voice TTS system known for its exceptional voice cloning capabilities.PythonGitHub stars
VALL-E-XPlachtaaOpen-source implementation of Microsoft's VALL-E X for zero-shot cross-lingual voice cloning.Python / PyTorchGitHub stars
PiperRhasspyFastest local TTS. Optimized for low-end hardware and offline use.C++ / PythonGitHub stars
AmphionAmphionOpen-source audio, music and speech generation toolkit containing multiple SOTA TTS models.PythonGitHub stars
VoiceCraftJason Li et al.Token-infilling neural codec model for zero-shot speech editing and TTS synthesis.Python / PyTorchGitHub stars
Kokoro-82MHexgradBest SOTA CPU TTS. Ultra-fast, studio quality, only 82M parameters.ONNX / PythonGitHub stars
StyleTTS 2yl4579Human-level TTS. Uses style diffusion, adversarial training, and large SLMs without phoneme duration models.Python / PyTorchGitHub stars
MeloTTSMyShellUltra-fast multilingual TTS running smoothly on CPU across English, Spanish, French, Chinese, Japanese, Korean.Python / PyTorchGitHub stars
sherpa-onnxNext-gen KaldiOffline multi-platform speech synthesis engine supporting VITS, Piper, and Kokoro on embedded/mobile/desktop.C++ / Python / GoGitHub stars
Parler-TTSHugging FaceLightweight, controllable speech generation with high naturalness.PythonHF
GLM-4-VoiceZhipu AIEnd-to-end voice model supporting real-time speech generation, emotion alteration, and bilingual dialogue.Python / PyTorchGitHub stars
Matcha-TTSShivam MehtaFast TTS architecture employing conditional flow matching, producing highly natural output.PythonGitHub stars
LocalModeLocalModeIn-browser TTS. Runs Kokoro (29 voices) and other AI models 100% in the browser via WebGPU/WASM. No server, no API keys, offline after first load.JavaScript / TypeScriptGitHub stars
VocelloPowerBeefNative Mac & iPhone app. Qwen3-TTS with preset speakers, natural-language voice design, and voice cloning. Runs entirely on Apple Silicon with no Python runtime, faster than realtime on an 8 GB M2.Swift / MLXGitHub stars
loudkitLoudReaderOn-device TTS with native SDKs. 28 voices in 10 languages, voice cloning from about 10 s of audio, and a local server with an OpenAI-compatible speech endpoint. PyTorch, ONNX Runtime and CoreML backends; the Swift, Go, Rust and TypeScript ports run without Python. Apache-2.0, derived from Chatterbox.Python / Swift / Go / Rust / TypeScriptGitHub stars

Advanced Voice Cloning & Neural Voice Synthesis 🧬

Dedicated resources and examples focusing on the latest in voice replication and advanced synthetic voice generation.

  • XTTS-v2 by Coqui: A breakthrough in voice cloning, capable of replicating a voice from just a 6-second audio clip, preserving emotion and speaking style.
  • GPT-SoVITS: Powerful few-shot voice cloning requiring only 5 seconds of sample audio to fine-tune or zero-shot clone with cross-lingual support.
  • CosyVoice 2 by Alibaba: SOTA multi-lingual voice cloning and emotional style transfer with sub-second prompt audio and dialect support.
  • Resemble AI's Chatterbox: Offers advanced zero-shot voice cloning capabilities, enabling instant voice replication without extensive training data.
  • ElevenLabs Voice Cloning: Provides robust tools for creating highly realistic voice clones, suitable for personalized audio content.
  • StyleTTS 2: State-of-the-art style diffusion TTS that matches human speech naturalness via continuous style vectors and adversarial training.
  • VoiceCraft: State-of-the-art token-infilling neural codec model capable of both zero-shot voice cloning and precise in-audio speech editing.
  • Suno Bark: A transformer-based text-to-audio model that generates highly naturalistic, multilingual speech, music, and sound effects. It excels at expressive speech with nuances like laughter, sighs, and crying.
  • MeloTTS: A multi-language, multi-speaker Text-to-Speech model capable of generating high-quality audio.

Hugging Face 🤗 - The Hub for TTS Models

Hugging Face has emerged as a central ecosystem for sharing, discovering, and experimenting with a vast array of pretrained Text-to-Speech models. Explore their extensive collection for diverse applications and research.

Notable Research Papers & Community Discussions 📝

Stay updated with the latest breakthroughs and discussions in the TTS community.

Exemplary Code Samples & Project Demos 💻

A collection of influential code repositories and product demonstrations showcasing various Text-to-Speech implementations and their output quality. Sorted by Year of Launch (Descending).

Project/SamplesPretrained ModelsCode LinkPaper/Arxiv IDOutput QualityYear of LaunchDescription
Gradbot Demos--CodeCodebaseA2026Eight voice agent demos (banking, hotel booking, 3D game NPCs) built on Gradium's real-time TTS/STT APIs.
Fish Speech v1.5--CodeCodebaseA+2026SOTA multilingual, multi-speaker model with superior naturalness.
GLM-4-Voice Samples--CodeCodebaseA2025End-to-end voice model with expressive bilingual dialogue.
Kokoro-82M Samples--Code--A2025Ultra-efficient CPU-based model with studio-quality output.
ChatTTS Samples--Code2406.03807A+2024Conversational dialogue speech synthesis with prosodic laughs and pauses.
CosyVoice Samples--Code2407.05407A+2024Multilingual voice generation with emotional control and multi-dialect support.
F5-TTS Samples--Code2410.06885A2024Diffusion-based zero-shot cloning with impressive prosody.
GPT-SoVITS Samples--CodeCodebaseA+2024Powerful few-shot voice cloning requiring only 5 seconds of sample audio.
MaskGCT Samples--Code2409.00750A2024Non-autoregressive zero-shot TTS using masked generative codec transformers.
MeloTTS Samples--CodeCodebaseB2024Multilingual, multi-speaker TTS model for high-quality audio generation.
Parler-TTS Samples--Code2402.01912B2024Samples from a lightweight model producing natural-sounding speech.
VoiceCraft Samples--Code2403.16973A2024Zero-shot speech editing and neural synthesis with token infilling.
Bark Samples (Suno.ai)--Code--A2023Samples from Suno's expressive text-to-audio model, including non-speech sounds.
StyleTTS 2 Samples--Code2306.07691A2023Human-level TTS with style diffusion and adversarial training.
VALL-E X Samples--Code2303.03926A2023Cross-lingual zero-shot speech synthesis and voice cloning.
XTTS-v2 Samples--Code2309.02055A2023Demonstrations of Coqui's advanced voice cloning with emotion transfer.
rayhane's Tacotron2 Samples------D2019Audio samples from an early Tacotron 2 implementation.
Google Tacotron + Style Transfer Sample (Official)----1803.09047A2018Official samples showcasing prosody and style transfer with Tacotron.
Kyubyong's DC-TTS on Nick Dataset Samples------D2018DC-TTS samples generated from the Nick dataset.
Kyubyong's Expressive Tacotron Samples--Code1803.09047D2018Samples demonstrating expressive speech synthesis with Tacotron.
Kyubyong's Tacotron on LJ Dataset SamplesDownload model----D2018Audio generated from Tacotron trained on the LJSpeech dataset.
Kyubyong's Tacotron on Nick Dataset Samples------D2018Tacotron samples from the Nick dataset.
Kyubyong's Tacotron on Web Dataset SamplesDownload model----D2018Tacotron speech output from the Web dataset.
mazzzystar's Tacotron-WaveRNN SamplesGet ModelCode--A2018Demonstrations from a Tacotron and WaveRNN hybrid model.
NVIDIA's Tacotron2 + WaveGlow SamplesDownload ModelCode--A2018Combined high-quality speech synthesis from Tacotron 2 and WaveGlow.
NVIDIA's WaveGlow SamplesDownload ModelCode1811.00002A2018High-fidelity audio generated by NVIDIA's WaveGlow vocoder.
syang1993's Tacotron + Style Transfer SamplesModel ErnstTmp (232k iter)--1803.09047 and 1803.09017C2018Samples demonstrating Tacotron with global style tokens for voice style transfer.
andabi's Deep Voice Conversion------D2017Demonstrations of deep voice conversion techniques.
Baidu's Deep Voice Samples (Official)------D2017Official audio demonstrations from Baidu's Deep Voice project.
Baidu's Deep Voice 3 Samples (Official)----1710.07654B2017Official samples from Deep Voice 3, showcasing advanced speech synthesis.
DeepMind Neural Discrete Representation Learning Samples (Official)----1711.00937B2017Samples demonstrating speech generated using VQ-VAE for neural discrete representation learning.
dhgrs's Implementation of Neural Discrete Representation Learning SamplesDownload ModelCode1711.00937D2017Audio generated using a Chainer implementation of VQ-VAE for speech.
Facebook Loop Samples (Official)Get model----D2017Official audio samples from Facebook's Loop project.
Google Tacotron2 Samples (Official)----1712.05884A2017Official, high-quality audio samples from the groundbreaking Tacotron 2 model.
keithito's Tacotron SamplesGet model----D2017Audio samples from keithito's Tacotron implementation.
Kyubyong's DC-TTS Kate Samples------D2017DC-TTS samples featuring the "Kate" voice.
Kyubyong's DC-TTS on LJ Dataset SamplesGet model----D2017DC-TTS generated speech from the LJSpeech dataset.
mazzzystar's RandomCNN Voice Transfer----1712.08363D2017Speech conversion samples using Random CNNs.
r9y9's Wavenet Vocoder Tacotron2 SamplesDownload Tacotron2 model - Download Wavenet model - Get models--1712.05884 and 1611.09482B2017Samples from a Tacotron 2 and WaveNet vocoder combination.
Griffin-Lim Samples------A1984Classic samples from the Griffin-Lim algorithm for spectrogram inversion.

Work in Progress & Future of Text-to-Speech 🚧

Ongoing projects and cutting-edge research shaping the next generation of AI voice synthesis.

If I missed your output sample/demo in this consolidation, just add and send a pull request. I will be more than happy to add it. Thanks!

Codelabs & Interactive Tutorials 🧪

Practical guides and interactive notebooks for experimenting with Text-to-Speech models.

Product Demos & Showcase Videos 🎥

Visual demonstrations of advanced Text-to-Speech and voice cloning in action.

Broader projects and research efforts that contribute to the Text-to-Speech ecosystem.

Arxiv Sanity Preserver - Key Papers in Speech Synthesis 📄

Explore influential academic papers and preprints in the field of Text-to-Speech and voice AI.

Star History

💬 Community & Support for Text-to-Speech Enthusiasts

Connect with the community, get support, and stay informed about the latest in TTS.

  • 📚 Documentation: Check out our official documentation for detailed guides and tutorials on utilizing TTS technologies.
  • 🗣️ Forum: Join our community forum to ask questions, share your Text-to-Speech projects, and connect with other users and developers.
  • 💬 Discord: Chat with us on Discord for real-time support and discussions on AI voice generation.
  • 🐦 Twitter: Follow us on Twitter for the latest news, updates, and insights into the world of synthetic speech.
  • 🐦 Github: Follow me on Github for the latest commits and updates on this and other AI projects.

🎯 Key Use Cases for AI Voice Generation

Explore how Text-to-Speech and Voice Cloning are being used across industries:

  • 🎙️ Podcast Automation: Convert written articles into high-quality audio episodes instantly.
  • 🎮 Video Game Development: Dynamic NPC dialogue using local-first TTS like Piper or Kokoro.
  • 🛠️ Customer Support: Low-latency conversational AI for 24/7 automated support.
  • 📖 Accessible E-Learning: Making educational content accessible with natural-sounding voices.
  • 🎬 Content Localization: Dubbing videos into multiple languages while preserving the original speaker's emotion.

❓ Frequently Asked Questions (FAQ) & SEO Insights

What is the best open-source Text-to-Speech model in 2026?

As of 2026, Kokoro-82M is widely considered the best for CPU-based local inference due to its studio quality and small footprint. For high-fidelity and expressive speech, F5-TTS and Fish Speech are leading the way in naturalness.

Are there free ElevenLabs alternatives for voice cloning?

Yes! Projects like Coqui XTTS-v2, GPT-SoVITS, and OpenVoice offer high-quality voice cloning for free. If you are looking for local-first alternatives, check out F5-TTS and CosyVoice.

How do I achieve low-latency TTS for AI agents?

To achieve sub-200ms latency, it is recommended to use Deepgram Aura, Cartesia Sonic, Smallest.ai, or optimized local models like Piper (C++ implementation) and Kokoro-82M with ONNX runtime.

Can I use these TTS models for commercial projects?

Many models listed here (like OpenAI TTS, ElevenLabs, and Azure AI Speech) have clear commercial tiers. For open-source models, look for those with MIT or Apache 2.0 licenses, such as Piper and Kokoro.


💖 Support & Sponsorship

If you find this collection of Text-to-Speech resources helpful, or if it has saved you time and effort in your AI voice generation endeavors, please consider sponsoring the development. Your support helps maintain the project, add new cutting-edge models and tools, and keep this initiative open-source and accessible to everyone.

Sponsor @ishandutta2007 on GitHub

Every contribution, no matter how small, makes a huge difference in advancing the Text-to-Speech landscape! 🙏

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

awesome-list
curated-list
deep-learning
nlp
prosody
style-tokens
style-transfer
tacotron
tacotron-2
tts
voice-cloning

Contributors

ishandutta2007

101 commits

AALG123

3 commits

easwee

1 commits

ishandutta2007/Awesome-Text-to-Speech

🎤 A curated list of the latest and most influential tools, models, and resources in the Text-to-Speech sector. 🌟 Star if you like it! 🌟

194

119 commits

updated Sep 14, 2026

See the code

README

Awesome Text-to-Speech Banner

Awesome Text-to-Speech (TTS) 🗣️: Best AI Voice Generation Models & Tools 2026

AwesomeDiscord Awesome GitHub stars GitHub forks GitHub license Python Version PRs Welcome GitHub Sponsors Last Commit Contributors Follow on Twitter GitHub followers


🚀 The Ultimate Guide to AI Voice Generation, Speech Synthesis, and Neural Voice Cloning

Welcome to the most comprehensive, meticulously curated, and continuously updated list of Text-to-Speech (TTS) resources. Whether you are looking for the best open-source TTS models of 2026, searching for low-latency TTS APIs for AI agents, or exploring high-fidelity voice cloning for content creation, you've found the right place.

[!TIP] Looking for the best ElevenLabs alternatives? This repository tracks the rapidly evolving landscape of both commercial SaaS and local-first neural speech synthesis.


🗺️ Quick Navigation


🔊 Why Explore Text-to-Speech?

Text-to-Speech technology has moved beyond robotic voices. Today, it powers:

  • Accessibility First: High-quality screen readers for the visually impaired.
  • Automated Content Creation: Realistic voiceovers for YouTube, podcasts, and e-learning.
  • Next-Gen AI Agents: Real-time conversational AI with human-like prosody.
  • Multilingual Support: Instant translation and dubbing for global reach.
  • Personalization: Custom voice clones for gaming and virtual assistants.

📈 Current State of Text-to-Speech (2026 Update)

The landscape of AI voice synthesis has shifted from basic concatenation to advanced Generative Speech Models. Key highlights:

  • Hyper-realistic and Natural Speech Synthesis: Innovations in deep learning and neural network architectures have led to highly natural, expressive, and emotionally nuanced synthetic voices. 🎤
  • Next-Generation Architectures: The adoption of State Space Models (SSMs), Diffusion Models, and advanced transformer-based architectures is offering superior performance, efficiency, and voice quality in speech generation. 🧠
  • Real-time Conversational AI: Significant advancements in reducing latency now enable real-time TTS, making conversational AI, virtual assistants, and live dubbing more natural and responsive. ⚡
  • Advanced Voice Cloning and Style Transfer: Cutting-edge techniques allow for high-fidelity voice cloning from minimal audio samples and the transfer of speaking style and emotion across different voices. 🎭
  • Multilingual and Cross-Lingual TTS: Models are increasingly capable of generating speech in numerous languages with accurate pronunciation and intonation, breaking down language barriers.

📚 Comprehensive List of Text-to-Speech (TTS) Resources 🌐

☁️ Cloud-based & Commercial AI Voice Generation Platforms

Leading platforms offering robust, scalable, and high-quality Text-to-Speech APIs and services for various applications. Sorted by Company Scale / Valuation (Descending).

🛠️ Service/Model🏢 Organization🌟 Key Features💰 Min. Monthly Subscription👥 Company Size🔗 Link
NVIDIA NeMoNVIDIAPlatform for building, training, and deploying generative AI models, including TTS and ASR.Free (API Credits) / Enterprise$3.3T+ (Market Cap)NVIDIA NeMo
Azure AI SpeechMicrosoftHigh-quality neural voices with advanced fine-tuning, emotion, and enterprise scalability.Free (0.5M chars/mo) / PAYG$3.2T+ (Market Cap)Azure AI Speech
Google Cloud TTSGooglePowerful TTS API with a large variety of natural-sounding voices and extensive customization.Free (1M+ chars/mo) / PAYG$2.2T+ (Market Cap)Google Cloud TTS
AWS PollyAmazonGenerative, Neural and Standard TTS voices with deep AWS ecosystem integration.Free (1M+ chars/mo) / PAYG$1.9T+ (Market Cap)AWS Polly
OpenAI TTSOpenAIHigh-quality, real-time streaming TTS models for applications requiring natural AI voices.Pay-as-you-go ($5 free credit)$852B (Valuation)OpenAI TTS
ElevenLabsElevenLabsState-of-the-art AI voice generator offering realistic voices, voice cloning, and AI dubbing.Free (10k chars/mo) / $5$11B (Valuation)ElevenLabs
SpeechifySpeechifyHighly popular consumer text-to-speech and developer Voice API with natural and premium voices.Free (Basic Voices) / $139/yr$1.5B (Valuation)Speechify
Deepgram AuraDeepgramSpecializing in low-latency TTS designed for real-time conversational AI and virtual interactions.Free ($200 credit) / PAYG$1.2B (Valuation)Deepgram Aura
Inworld AIInworld AICharacter-driven conversational voice engine and real-time speech generation for interactive agents.Free (Basic / 100k API credits) / $20$500M (Valuation)Inworld AI
Hume AI (EVI)Hume AIEmpathic Voice Interface with emotional prosody detection and low-latency expressive speech.Free ($10 credit) / PAYG ($0.036/min)$250M (Valuation)Hume AI
Cartesia SonicCartesiaSub-100ms ultra-low latency TTS designed for real-time AI agents.Free (API Credits) / PAYG$200M (Valuation)Cartesia
GradiumGradiumReal-time TTS and STT for voice agents. 158ms P50 time-to-first-audio, streaming and instant voice cloning.Free (45k credits) / $13$100M (Seed Raised)Gradium
Murf.aiMurf.aiAI voiceovers with a built-in video editor, ideal for creators and presentations.Free (10 mins total) / $19$46M (Valuation)Murf.ai
WellSaid LabsWellSaid LabsEnterprise AI voice platform with natural studio voices, fine phonetic control, and brand voice avatars.Free (7-day trial / 50 clips) / $49$40M (Valuation)WellSaid Labs
Resemble AIResemble AIHigh-fidelity voice cloning, neural watermarking, deepfake detection, and real-time speech synthesis.Free (Free trial) / $29$30M (Valuation)Resemble AI
LMNTLMNTLightning-fast TTS API with exceptional naturalness, great for interactive voice apps.Free (API Credits) / PAYG$21M (Estimated)LMNT
Lovo.ai (Genny)LOVONext-gen AI voice generator with 500+ voices in 100+ languages and granular pitch/emphasis controls.Free (14-day Pro trial) / $29$20M (Estimated)Lovo.ai
Play.htPlay.htProfessional AI voices and "Ultra-Realistic" studio editor for long-form content.Free (12.5k chars) / $39$15M (Valuation)Play.ht
Smallest.ai (Waves)Smallest.aiLightning-fast TTS API with sub-100ms latency designed for real-time conversational agents and voice bots.Free (10k chars/mo) / PAYG ($0.08/1k chars)$12M (Valuation)Smallest.ai
Rime LabsRimeUltra-fast expressive voice API built specifically for interactive conversational AI applications.Free ($5 credit) / PAYG$10M (Estimated)Rime
Soniox TTSSonioxReal-time streaming TTS API for conversational AI voice agents in 60+ languages with multilingual voices.Pay-as-you-go (~$0.70/hr)$10M (Estimated)Soniox
Neets.aiNeetsExtremely fast and affordable TTS APIs starting at $0.0004 per 1k characters.Free (API Credits) / PAYG$5M (Estimated)Neets.ai
GandrGandrTTS API for voice agents. Word error rate 1.98 percent against a 2.17 percent human reference on the same scorer, one voice in 23 languages, every render watermarked. Python/JS SDKs, LiveKit plugin, MCP server.From $10/mo. Flat-rate unmetered streams from $150/mo<$1M (Indie)Gandr
SpokioSpokioOffline macOS text-to-speech app with local voice cloning, batch export, and no cloud uploads.Free (API Credits) / Enterprise<$1M (Indie)Spokio
PHANTOM VOICESPHANTOM VOICES10 free professional AI voice clones via public REST API. Zero cost, commercial rights cleared. 29 platform configs (Vapi, Retell AI, etc). Multilingual (9+ languages). AI-powered recommendation.Free (API Credits) / Enterprise<$1M (Indie)PHANTOM VOICES
RunAPI ElevenLabs SDKRunAPIMulti-language SDKs for ElevenLabs text-to-speech, dialogue generation, sound effects, transcription, and audio isolation workflows.Pay-as-you-go<$1M (Indie)RunAPI ElevenLabs SDK
AudexumAudexumText-to-speech and speech-to-text in one API: 43 voices, 32 TTS languages, 25 STT languages. ElevenLabs-compatible endpoint, so switching is a base-URL change. EU-hosted, every output watermarked.Free (30k credits at signup, 3k/mo after) / EUR 4<$1M (Indie)Audexum

🏗️ Open-Source Text-to-Speech Libraries & Local-First Projects

If you are looking for free text-to-speech models for commercial use or want to run TTS locally on a CPU, these open-source projects provide the best balance of quality and privacy. Sorted by GitHub Star Counts (Descending).

Sound Wave Animation
🛠️ Service/Model🏢 Organization🌟 Key Features🗣️ Primary Language📁 Github_Repository
🐸 Coqui TTSCoqui (community)Supports 1100+ languages, zero-shot voice cloning, and fine-tuning. Note: Coqui AI (the company) shut down in late 2023; actively maintained by the community at idiap/coqui-ai-TTS.Python / MultilingualGitHub stars
GPT-SoVITSRVC-BossZero-shot & few-shot voice cloning. Requires only 5 seconds of sample audio for cross-lingual synthesis.Python / PyTorchGitHub stars
BarkSunoTransformer-based text-to-audio model capable of highly expressive speech, music, laughs, and sighs.Python / PyTorchGitHub stars
ChatTTS2noiseConversational text-to-speech model specially optimized for dialogue and natural conversational flow.PythonGitHub stars
OpenVoiceMyShellHighly versatile and instant voice cloning that requires only a short audio clip.PythonGitHub stars
Fish SpeechFish AudioSOTA multilingual, multi-speaker model with superior naturalness.PythonGitHub stars
ChatterboxResemble AIAdvanced neural voice synthesis with emotion control and high-fidelity cloning.PythonGitHub stars
CosyVoiceAlibabaExcellent multilingual and zero-shot voice cloning model capable of high fidelity.PythonGitHub stars
KittenTTSKittenMLONNX-based library for low-latency TTS without requiring a GPU.Python / ONNXGitHub stars
F5-TTSSWividFlow Matching TTS. Incredible naturalness and prosody using DiT architectures.PythonGitHub stars
Tortoise-TTSJames BetkerPowerful multi-voice TTS system known for its exceptional voice cloning capabilities.PythonGitHub stars
VALL-E-XPlachtaaOpen-source implementation of Microsoft's VALL-E X for zero-shot cross-lingual voice cloning.Python / PyTorchGitHub stars
PiperRhasspyFastest local TTS. Optimized for low-end hardware and offline use.C++ / PythonGitHub stars
AmphionAmphionOpen-source audio, music and speech generation toolkit containing multiple SOTA TTS models.PythonGitHub stars
VoiceCraftJason Li et al.Token-infilling neural codec model for zero-shot speech editing and TTS synthesis.Python / PyTorchGitHub stars
Kokoro-82MHexgradBest SOTA CPU TTS. Ultra-fast, studio quality, only 82M parameters.ONNX / PythonGitHub stars
StyleTTS 2yl4579Human-level TTS. Uses style diffusion, adversarial training, and large SLMs without phoneme duration models.Python / PyTorchGitHub stars
MeloTTSMyShellUltra-fast multilingual TTS running smoothly on CPU across English, Spanish, French, Chinese, Japanese, Korean.Python / PyTorchGitHub stars
sherpa-onnxNext-gen KaldiOffline multi-platform speech synthesis engine supporting VITS, Piper, and Kokoro on embedded/mobile/desktop.C++ / Python / GoGitHub stars
Parler-TTSHugging FaceLightweight, controllable speech generation with high naturalness.PythonHF
GLM-4-VoiceZhipu AIEnd-to-end voice model supporting real-time speech generation, emotion alteration, and bilingual dialogue.Python / PyTorchGitHub stars
Matcha-TTSShivam MehtaFast TTS architecture employing conditional flow matching, producing highly natural output.PythonGitHub stars
LocalModeLocalModeIn-browser TTS. Runs Kokoro (29 voices) and other AI models 100% in the browser via WebGPU/WASM. No server, no API keys, offline after first load.JavaScript / TypeScriptGitHub stars
VocelloPowerBeefNative Mac & iPhone app. Qwen3-TTS with preset speakers, natural-language voice design, and voice cloning. Runs entirely on Apple Silicon with no Python runtime, faster than realtime on an 8 GB M2.Swift / MLXGitHub stars
loudkitLoudReaderOn-device TTS with native SDKs. 28 voices in 10 languages, voice cloning from about 10 s of audio, and a local server with an OpenAI-compatible speech endpoint. PyTorch, ONNX Runtime and CoreML backends; the Swift, Go, Rust and TypeScript ports run without Python. Apache-2.0, derived from Chatterbox.Python / Swift / Go / Rust / TypeScriptGitHub stars

Advanced Voice Cloning & Neural Voice Synthesis 🧬

Dedicated resources and examples focusing on the latest in voice replication and advanced synthetic voice generation.

  • XTTS-v2 by Coqui: A breakthrough in voice cloning, capable of replicating a voice from just a 6-second audio clip, preserving emotion and speaking style.
  • GPT-SoVITS: Powerful few-shot voice cloning requiring only 5 seconds of sample audio to fine-tune or zero-shot clone with cross-lingual support.
  • CosyVoice 2 by Alibaba: SOTA multi-lingual voice cloning and emotional style transfer with sub-second prompt audio and dialect support.
  • Resemble AI's Chatterbox: Offers advanced zero-shot voice cloning capabilities, enabling instant voice replication without extensive training data.
  • ElevenLabs Voice Cloning: Provides robust tools for creating highly realistic voice clones, suitable for personalized audio content.
  • StyleTTS 2: State-of-the-art style diffusion TTS that matches human speech naturalness via continuous style vectors and adversarial training.
  • VoiceCraft: State-of-the-art token-infilling neural codec model capable of both zero-shot voice cloning and precise in-audio speech editing.
  • Suno Bark: A transformer-based text-to-audio model that generates highly naturalistic, multilingual speech, music, and sound effects. It excels at expressive speech with nuances like laughter, sighs, and crying.
  • MeloTTS: A multi-language, multi-speaker Text-to-Speech model capable of generating high-quality audio.

Hugging Face 🤗 - The Hub for TTS Models

Hugging Face has emerged as a central ecosystem for sharing, discovering, and experimenting with a vast array of pretrained Text-to-Speech models. Explore their extensive collection for diverse applications and research.

Notable Research Papers & Community Discussions 📝

Stay updated with the latest breakthroughs and discussions in the TTS community.

Exemplary Code Samples & Project Demos 💻

A collection of influential code repositories and product demonstrations showcasing various Text-to-Speech implementations and their output quality. Sorted by Year of Launch (Descending).

Project/SamplesPretrained ModelsCode LinkPaper/Arxiv IDOutput QualityYear of LaunchDescription
Gradbot Demos--CodeCodebaseA2026Eight voice agent demos (banking, hotel booking, 3D game NPCs) built on Gradium's real-time TTS/STT APIs.
Fish Speech v1.5--CodeCodebaseA+2026SOTA multilingual, multi-speaker model with superior naturalness.
GLM-4-Voice Samples--CodeCodebaseA2025End-to-end voice model with expressive bilingual dialogue.
Kokoro-82M Samples--Code--A2025Ultra-efficient CPU-based model with studio-quality output.
ChatTTS Samples--Code2406.03807A+2024Conversational dialogue speech synthesis with prosodic laughs and pauses.
CosyVoice Samples--Code2407.05407A+2024Multilingual voice generation with emotional control and multi-dialect support.
F5-TTS Samples--Code2410.06885A2024Diffusion-based zero-shot cloning with impressive prosody.
GPT-SoVITS Samples--CodeCodebaseA+2024Powerful few-shot voice cloning requiring only 5 seconds of sample audio.
MaskGCT Samples--Code2409.00750A2024Non-autoregressive zero-shot TTS using masked generative codec transformers.
MeloTTS Samples--CodeCodebaseB2024Multilingual, multi-speaker TTS model for high-quality audio generation.
Parler-TTS Samples--Code2402.01912B2024Samples from a lightweight model producing natural-sounding speech.
VoiceCraft Samples--Code2403.16973A2024Zero-shot speech editing and neural synthesis with token infilling.
Bark Samples (Suno.ai)--Code--A2023Samples from Suno's expressive text-to-audio model, including non-speech sounds.
StyleTTS 2 Samples--Code2306.07691A2023Human-level TTS with style diffusion and adversarial training.
VALL-E X Samples--Code2303.03926A2023Cross-lingual zero-shot speech synthesis and voice cloning.
XTTS-v2 Samples--Code2309.02055A2023Demonstrations of Coqui's advanced voice cloning with emotion transfer.
rayhane's Tacotron2 Samples------D2019Audio samples from an early Tacotron 2 implementation.
Google Tacotron + Style Transfer Sample (Official)----1803.09047A2018Official samples showcasing prosody and style transfer with Tacotron.
Kyubyong's DC-TTS on Nick Dataset Samples------D2018DC-TTS samples generated from the Nick dataset.
Kyubyong's Expressive Tacotron Samples--Code1803.09047D2018Samples demonstrating expressive speech synthesis with Tacotron.
Kyubyong's Tacotron on LJ Dataset SamplesDownload model----D2018Audio generated from Tacotron trained on the LJSpeech dataset.
Kyubyong's Tacotron on Nick Dataset Samples------D2018Tacotron samples from the Nick dataset.
Kyubyong's Tacotron on Web Dataset SamplesDownload model----D2018Tacotron speech output from the Web dataset.
mazzzystar's Tacotron-WaveRNN SamplesGet ModelCode--A2018Demonstrations from a Tacotron and WaveRNN hybrid model.
NVIDIA's Tacotron2 + WaveGlow SamplesDownload ModelCode--A2018Combined high-quality speech synthesis from Tacotron 2 and WaveGlow.
NVIDIA's WaveGlow SamplesDownload ModelCode1811.00002A2018High-fidelity audio generated by NVIDIA's WaveGlow vocoder.
syang1993's Tacotron + Style Transfer SamplesModel ErnstTmp (232k iter)--1803.09047 and 1803.09017C2018Samples demonstrating Tacotron with global style tokens for voice style transfer.
andabi's Deep Voice Conversion------D2017Demonstrations of deep voice conversion techniques.
Baidu's Deep Voice Samples (Official)------D2017Official audio demonstrations from Baidu's Deep Voice project.
Baidu's Deep Voice 3 Samples (Official)----1710.07654B2017Official samples from Deep Voice 3, showcasing advanced speech synthesis.
DeepMind Neural Discrete Representation Learning Samples (Official)----1711.00937B2017Samples demonstrating speech generated using VQ-VAE for neural discrete representation learning.
dhgrs's Implementation of Neural Discrete Representation Learning SamplesDownload ModelCode1711.00937D2017Audio generated using a Chainer implementation of VQ-VAE for speech.
Facebook Loop Samples (Official)Get model----D2017Official audio samples from Facebook's Loop project.
Google Tacotron2 Samples (Official)----1712.05884A2017Official, high-quality audio samples from the groundbreaking Tacotron 2 model.
keithito's Tacotron SamplesGet model----D2017Audio samples from keithito's Tacotron implementation.
Kyubyong's DC-TTS Kate Samples------D2017DC-TTS samples featuring the "Kate" voice.
Kyubyong's DC-TTS on LJ Dataset SamplesGet model----D2017DC-TTS generated speech from the LJSpeech dataset.
mazzzystar's RandomCNN Voice Transfer----1712.08363D2017Speech conversion samples using Random CNNs.
r9y9's Wavenet Vocoder Tacotron2 SamplesDownload Tacotron2 model - Download Wavenet model - Get models--1712.05884 and 1611.09482B2017Samples from a Tacotron 2 and WaveNet vocoder combination.
Griffin-Lim Samples------A1984Classic samples from the Griffin-Lim algorithm for spectrogram inversion.

Work in Progress & Future of Text-to-Speech 🚧

Ongoing projects and cutting-edge research shaping the next generation of AI voice synthesis.

If I missed your output sample/demo in this consolidation, just add and send a pull request. I will be more than happy to add it. Thanks!

Codelabs & Interactive Tutorials 🧪

Practical guides and interactive notebooks for experimenting with Text-to-Speech models.

Product Demos & Showcase Videos 🎥

Visual demonstrations of advanced Text-to-Speech and voice cloning in action.

Broader projects and research efforts that contribute to the Text-to-Speech ecosystem.

Arxiv Sanity Preserver - Key Papers in Speech Synthesis 📄

Explore influential academic papers and preprints in the field of Text-to-Speech and voice AI.

Star History

💬 Community & Support for Text-to-Speech Enthusiasts

Connect with the community, get support, and stay informed about the latest in TTS.

  • 📚 Documentation: Check out our official documentation for detailed guides and tutorials on utilizing TTS technologies.
  • 🗣️ Forum: Join our community forum to ask questions, share your Text-to-Speech projects, and connect with other users and developers.
  • 💬 Discord: Chat with us on Discord for real-time support and discussions on AI voice generation.
  • 🐦 Twitter: Follow us on Twitter for the latest news, updates, and insights into the world of synthetic speech.
  • 🐦 Github: Follow me on Github for the latest commits and updates on this and other AI projects.

🎯 Key Use Cases for AI Voice Generation

Explore how Text-to-Speech and Voice Cloning are being used across industries:

  • 🎙️ Podcast Automation: Convert written articles into high-quality audio episodes instantly.
  • 🎮 Video Game Development: Dynamic NPC dialogue using local-first TTS like Piper or Kokoro.
  • 🛠️ Customer Support: Low-latency conversational AI for 24/7 automated support.
  • 📖 Accessible E-Learning: Making educational content accessible with natural-sounding voices.
  • 🎬 Content Localization: Dubbing videos into multiple languages while preserving the original speaker's emotion.

❓ Frequently Asked Questions (FAQ) & SEO Insights

What is the best open-source Text-to-Speech model in 2026?

As of 2026, Kokoro-82M is widely considered the best for CPU-based local inference due to its studio quality and small footprint. For high-fidelity and expressive speech, F5-TTS and Fish Speech are leading the way in naturalness.

Are there free ElevenLabs alternatives for voice cloning?

Yes! Projects like Coqui XTTS-v2, GPT-SoVITS, and OpenVoice offer high-quality voice cloning for free. If you are looking for local-first alternatives, check out F5-TTS and CosyVoice.

How do I achieve low-latency TTS for AI agents?

To achieve sub-200ms latency, it is recommended to use Deepgram Aura, Cartesia Sonic, Smallest.ai, or optimized local models like Piper (C++ implementation) and Kokoro-82M with ONNX runtime.

Can I use these TTS models for commercial projects?

Many models listed here (like OpenAI TTS, ElevenLabs, and Azure AI Speech) have clear commercial tiers. For open-source models, look for those with MIT or Apache 2.0 licenses, such as Piper and Kokoro.


💖 Support & Sponsorship

If you find this collection of Text-to-Speech resources helpful, or if it has saved you time and effort in your AI voice generation endeavors, please consider sponsoring the development. Your support helps maintain the project, add new cutting-edge models and tools, and keep this initiative open-source and accessible to everyone.

Sponsor @ishandutta2007 on GitHub

Every contribution, no matter how small, makes a huge difference in advancing the Text-to-Speech landscape! 🙏

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

awesome-list
curated-list
deep-learning
nlp
prosody
style-tokens
style-transfer
tacotron
tacotron-2
tts
voice-cloning

Contributors

ishandutta2007

101 commits

AALG123

3 commits

easwee

1 commits