🎙️ VoxSherpa TTS Offline Neural Text-to-Speech Engine for Android ⚡ Sherpa-ONNX powered 🔊 Natural voice synthesis 📱 Fully offline processing 🚀 No cloud • No limits
See the code
🛡️ Privacy-focused user? Please check our Documentation before downloading.
VoxSherpa TTS is listed in the official README of k2-fsa/sherpa-onnx — the core inference library powering this app.
Most TTS apps make you choose between quality and privacy. Cloud-based tools like ElevenLabs sound incredible — but they require internet, send your text to remote servers, and charge per character.
VoxSherpa breaks that tradeoff.
It runs two professional-grade neural engines entirely on your device:
| Engine | Quality | Speed | Best For |
|---|---|---|---|
| 🧠 Kokoro-82M | Studio-grade · rivals ElevenLabs | Slower on budget hardware | Audiobooks, voiceovers, professional content |
| ⚡ Piper / VITS | Natural · clear · multi-speaker | Fast on any device | Daily use, dialogue synthesis, quick synthesis |
| Generate | Models | Library | Settings |
|---|---|---|---|
![]() | ![]() | ![]() | ![]() |
[speaker] tag system to assign each line to a distinct voice directly inside your script:
[speaker:1] Hello, how are you?
[speaker:2] I'm good, thanks! How about you?
.onnx models from local storage[whisper], [angry], [happy] supportUser Text
│
├─── Kokoro Engine (KokoroEngine.java)
│ └── Sherpa-ONNX JNI → ONNX Runtime → CPU/NNAPI
│ └── kokoro-multi-lang-v1_0
│
└─── Piper / VITS Engine (VoiceEngine.java)
└── Sherpa-ONNX JNI → ONNX Runtime → CPU
└── VITS model (language-specific)
Built with:
Generation speed depends entirely on your device's processor:
| Device Tier | Kokoro | Piper |
|---|---|---|
| 🟢 Flagship (Snapdragon 8 Gen 3) | ~20–40 sec/min audio | ~5 sec/min audio |
| 🟡 Mid-range (8-core) | ~60–90 sec/min audio | ~10 sec/min audio |
| 🔴 Budget (6-core) | ~2–3 min/min audio | ~20 sec/min audio |
Kokoro prioritizes quality over speed by design. It uses the same 82M parameter architecture that powers premium commercial TTS — running it entirely offline on a mobile CPU is genuinely pushing the hardware limits.
Requirements: Android 11+ · ARM64 · ~500 MB free storage recommended (for models)
VoxSherpa supports importing custom .onnx models without any server:
.onnx model + tokens.txt on device storageCompatible with any Sherpa-ONNX compatible TTS model.
VoxSherpa is open source. Contributions welcome:
Copyright (C) 2025 CodeBySonu95
This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.
https://www.gnu.org/licenses/gpl-3.0.html
Built with obsession. Runs without internet.
VoxSherpa — Because your voice deserves to stay yours.
43 followers · starred Jun 2026
108 followers · starred May 2026
115 followers · starred May 2026
Java
94.3%
HTML
5.7%
🎙️ VoxSherpa TTS Offline Neural Text-to-Speech Engine for Android ⚡ Sherpa-ONNX powered 🔊 Natural voice synthesis 📱 Fully offline processing 🚀 No cloud • No limits
See the code
🛡️ Privacy-focused user? Please check our Documentation before downloading.
VoxSherpa TTS is listed in the official README of k2-fsa/sherpa-onnx — the core inference library powering this app.
Most TTS apps make you choose between quality and privacy. Cloud-based tools like ElevenLabs sound incredible — but they require internet, send your text to remote servers, and charge per character.
VoxSherpa breaks that tradeoff.
It runs two professional-grade neural engines entirely on your device:
| Engine | Quality | Speed | Best For |
|---|---|---|---|
| 🧠 Kokoro-82M | Studio-grade · rivals ElevenLabs | Slower on budget hardware | Audiobooks, voiceovers, professional content |
| ⚡ Piper / VITS | Natural · clear · multi-speaker | Fast on any device | Daily use, dialogue synthesis, quick synthesis |
| Generate | Models | Library | Settings |
|---|---|---|---|
![]() | ![]() | ![]() | ![]() |
[speaker] tag system to assign each line to a distinct voice directly inside your script:
[speaker:1] Hello, how are you?
[speaker:2] I'm good, thanks! How about you?
.onnx models from local storage[whisper], [angry], [happy] supportUser Text
│
├─── Kokoro Engine (KokoroEngine.java)
│ └── Sherpa-ONNX JNI → ONNX Runtime → CPU/NNAPI
│ └── kokoro-multi-lang-v1_0
│
└─── Piper / VITS Engine (VoiceEngine.java)
└── Sherpa-ONNX JNI → ONNX Runtime → CPU
└── VITS model (language-specific)
Built with:
Generation speed depends entirely on your device's processor:
| Device Tier | Kokoro | Piper |
|---|---|---|
| 🟢 Flagship (Snapdragon 8 Gen 3) | ~20–40 sec/min audio | ~5 sec/min audio |
| 🟡 Mid-range (8-core) | ~60–90 sec/min audio | ~10 sec/min audio |
| 🔴 Budget (6-core) | ~2–3 min/min audio | ~20 sec/min audio |
Kokoro prioritizes quality over speed by design. It uses the same 82M parameter architecture that powers premium commercial TTS — running it entirely offline on a mobile CPU is genuinely pushing the hardware limits.
Requirements: Android 11+ · ARM64 · ~500 MB free storage recommended (for models)
VoxSherpa supports importing custom .onnx models without any server:
.onnx model + tokens.txt on device storageCompatible with any Sherpa-ONNX compatible TTS model.
VoxSherpa is open source. Contributions welcome:
Copyright (C) 2025 CodeBySonu95
This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.
https://www.gnu.org/licenses/gpl-3.0.html
Built with obsession. Runs without internet.
VoxSherpa — Because your voice deserves to stay yours.
43 followers · starred Jun 2026
108 followers · starred May 2026
115 followers · starred May 2026
Java
94.3%
HTML
5.7%