ayousanz/piper-plus-demo

Space

piper-plus Demo

4

79 commits

4 linked in READMEs

updated Sep 30, 2026

See the code

README

piper-plus Demo

A web-based demo for piper-plus, featuring high-quality text-to-speech synthesis. The code supports 8 languages (Japanese, English, Chinese, Korean, Spanish, French, Portuguese, Swedish); the current demo model covers 6 languages (Korean and Swedish models not yet trained).

Features

  • πŸ‡―πŸ‡΅ Japanese TTS: High-quality Japanese speech synthesis using OpenJTalk phonemization
  • πŸ‡ΊπŸ‡Έ English TTS: Natural English speech synthesis
  • πŸ‡¨πŸ‡³ Chinese TTS: Mandarin Chinese speech synthesis using pypinyin
  • πŸ‡ͺπŸ‡Έ Spanish TTS: Spanish speech synthesis with rule-based phonemization
  • πŸ‡«πŸ‡· French TTS: French speech synthesis with rule-based phonemization
  • πŸ‡΅πŸ‡Ή Portuguese TTS: Portuguese speech synthesis with rule-based phonemization
  • πŸš€ Fast Inference: ONNX Runtime for efficient CPU-based inference
  • πŸŽ›οΈ Adjustable Parameters: Control speech speed, expressiveness, and phoneme duration
  • 🌐 Web Interface: Easy-to-use Gradio interface

Supported Languages

CodeLanguageScriptPhonemizerIn demo model
jaJapaneseHiragana/Katakana/Kanjipyopenjtalkβœ…
enEnglishLating2p-enβœ…
zhChinese (Mandarin)Simplified Chinesepypinyinβœ…
koKoreanHangulg2pk2❌ (G2P only β€” model not trained yet)
esSpanishLatinRule-basedβœ…
frFrenchLatinRule-basedβœ…
ptPortugueseLatinRule-basedβœ…
svSwedishLatinRule-based❌ (G2P only β€” model not trained yet)

Models

This demo uses a multilingual model trained on 6 languages with 571 speakers and 508,187 utterances:

  • Multilingual 6-lang (Medium): Supports Japanese, English, Chinese, Spanish, French, and Portuguese with a unified 173-symbol model

Usage

  1. Select a model from the dropdown
  2. Select the language of your input text
  3. Enter your text in the input field (see sample texts below)
  4. Adjust advanced settings if needed
  5. Click "Generate Speech" to synthesize

Sample Texts by Language

LanguageSample Text
Japanese (ja)γ“γ‚“γ«γ‘γ―γ€δ»Šζ—₯は良い倩気ですね。
English (en)Hello, how are you today?
Chinese (zh)δ½ ε₯½οΌŒδ»Šε€©ε€©ζ°”εΎˆε₯½γ€‚
Spanish (es)Hola, como estas hoy?
French (fr)Bonjour, comment allez-vous?
Portuguese (pt)Ola, como voce esta hoje?

Code-Switching (Mixed Language)

The multilingual model supports code-switching within a single utterance. The language detector automatically identifies language segments:

今ζ—₯はgood morningですね  β†’  JA + EN mixed

Technical Details

  • Framework: ONNX Runtime (CPU inference)
  • Phonemization:
    • Japanese: pyopenjtalk (OpenJTalk-based)
    • English: g2p-en (Apache-2.0, espeak-ng compatible)
    • Chinese: pypinyin (MIT)
    • Spanish / French / Portuguese: Rule-based (no external dependency)
  • Language Detection: Unicode range-based automatic detection
  • Audio: 16-bit PCM WAV output
  • Model Architecture: VITS with language embedding (emb_lang), 173 phoneme symbols, 6 language IDs

Local Development

# Clone the repository
git clone https://github.com/ayutaz/piper-plus.git
cd piper-plus/huggingface-space

# Install requirements
uv pip install -r requirements.txt

# Run the app
python app.py

Credits

  • Piper TTS by Rhasspy
  • Multilingual (6-language) enhancements by ayutaz
  • Chinese phonemization: pypinyin (MIT)
  • English G2P: g2p-en (Apache-2.0)

License

This project is licensed under the MIT License. See the original Piper repository for more details.


Last updated: 2026-03-18 - 6-language multilingual support (ja/en/zh/es/fr/pt)

gradio

ayousanz/piper-plus-demo

Space

piper-plus Demo

4

79 commits

4 linked in READMEs

updated Sep 30, 2026

See the code

README

piper-plus Demo

A web-based demo for piper-plus, featuring high-quality text-to-speech synthesis. The code supports 8 languages (Japanese, English, Chinese, Korean, Spanish, French, Portuguese, Swedish); the current demo model covers 6 languages (Korean and Swedish models not yet trained).

Features

  • πŸ‡―πŸ‡΅ Japanese TTS: High-quality Japanese speech synthesis using OpenJTalk phonemization
  • πŸ‡ΊπŸ‡Έ English TTS: Natural English speech synthesis
  • πŸ‡¨πŸ‡³ Chinese TTS: Mandarin Chinese speech synthesis using pypinyin
  • πŸ‡ͺπŸ‡Έ Spanish TTS: Spanish speech synthesis with rule-based phonemization
  • πŸ‡«πŸ‡· French TTS: French speech synthesis with rule-based phonemization
  • πŸ‡΅πŸ‡Ή Portuguese TTS: Portuguese speech synthesis with rule-based phonemization
  • πŸš€ Fast Inference: ONNX Runtime for efficient CPU-based inference
  • πŸŽ›οΈ Adjustable Parameters: Control speech speed, expressiveness, and phoneme duration
  • 🌐 Web Interface: Easy-to-use Gradio interface

Supported Languages

CodeLanguageScriptPhonemizerIn demo model
jaJapaneseHiragana/Katakana/Kanjipyopenjtalkβœ…
enEnglishLating2p-enβœ…
zhChinese (Mandarin)Simplified Chinesepypinyinβœ…
koKoreanHangulg2pk2❌ (G2P only β€” model not trained yet)
esSpanishLatinRule-basedβœ…
frFrenchLatinRule-basedβœ…
ptPortugueseLatinRule-basedβœ…
svSwedishLatinRule-based❌ (G2P only β€” model not trained yet)

Models

This demo uses a multilingual model trained on 6 languages with 571 speakers and 508,187 utterances:

  • Multilingual 6-lang (Medium): Supports Japanese, English, Chinese, Spanish, French, and Portuguese with a unified 173-symbol model

Usage

  1. Select a model from the dropdown
  2. Select the language of your input text
  3. Enter your text in the input field (see sample texts below)
  4. Adjust advanced settings if needed
  5. Click "Generate Speech" to synthesize

Sample Texts by Language

LanguageSample Text
Japanese (ja)γ“γ‚“γ«γ‘γ―γ€δ»Šζ—₯は良い倩気ですね。
English (en)Hello, how are you today?
Chinese (zh)δ½ ε₯½οΌŒδ»Šε€©ε€©ζ°”εΎˆε₯½γ€‚
Spanish (es)Hola, como estas hoy?
French (fr)Bonjour, comment allez-vous?
Portuguese (pt)Ola, como voce esta hoje?

Code-Switching (Mixed Language)

The multilingual model supports code-switching within a single utterance. The language detector automatically identifies language segments:

今ζ—₯はgood morningですね  β†’  JA + EN mixed

Technical Details

  • Framework: ONNX Runtime (CPU inference)
  • Phonemization:
    • Japanese: pyopenjtalk (OpenJTalk-based)
    • English: g2p-en (Apache-2.0, espeak-ng compatible)
    • Chinese: pypinyin (MIT)
    • Spanish / French / Portuguese: Rule-based (no external dependency)
  • Language Detection: Unicode range-based automatic detection
  • Audio: 16-bit PCM WAV output
  • Model Architecture: VITS with language embedding (emb_lang), 173 phoneme symbols, 6 language IDs

Local Development

# Clone the repository
git clone https://github.com/ayutaz/piper-plus.git
cd piper-plus/huggingface-space

# Install requirements
uv pip install -r requirements.txt

# Run the app
python app.py

Credits

  • Piper TTS by Rhasspy
  • Multilingual (6-language) enhancements by ayutaz
  • Chinese phonemization: pypinyin (MIT)
  • English G2P: g2p-en (Apache-2.0)

License

This project is licensed under the MIT License. See the original Piper repository for more details.


Last updated: 2026-03-18 - 6-language multilingual support (ja/en/zh/es/fr/pt)

gradio