vvolhejn/pocket-tts-czech

Model

0

stars

4

commits

1

linked in READMEs

Aug 25, 2026

updated

czech
pocket-tts
safetensors
text-to-speech
tts

README

Pocket TTS — Czech (6 layers)

A Czech voice-cloning TTS model for Pocket TTS. Six transformer layers, runs faster than realtime on CPU.

Usage

uvx pocket-tts generate \
  --config hf://vvolhejn/pocket-tts-czech/czech.yaml@7c1fbd0acba765617749dd17f3dbddc2be791cc7 \
  --voice your_voice.wav \
  --text "Dobrý den, toto je český model."

The voice must be an audio file (or a state exported with pocket-tts export-voice from these weights). The named catalog voices — alba, cosette, … — are conditioning states precomputed with the released English weights and will not work here.

Sample voices cut from held-out ParCzech speakers are in voices/:

uvx pocket-tts generate \
  --config hf://vvolhejn/pocket-tts-czech/czech.yaml@7c1fbd0acba765617749dd17f3dbddc2be791cc7 \
  --voice https://huggingface.co/vvolhejn/pocket-tts-czech/resolve/main/voices/cs_m_zenisek.wav \
  --text "Dobrý den, toto je český model."

Training

corpusParCzech4Speech (Czech parliamentary speech, CC-BY 4.0)
training data976 h, 547,597 utterances, 524 speakers
tokenizersentencepiece fitted on the ParCzech transcripts, vocab 3999
teacher24 layers, LSD from scratch, 250k steps, lr 2e-4 constant, flow_batch_multiplier 4
student6 layers, depth-distilled from the teacher's EMA weights, 100k steps, lr 4e-4 cosine, distill_cfg_coef 2.0
weightsEMA (decay 0.9999)
recommended--temperature 0.3, cfg 1 (the default; guidance is baked into the student)

Trained with the Pocket TTS training code, following the non-English recipe.

Evaluation

The Pocket TTS eval pipeline (WER / speaker similarity / UTMOS) is English-only, so no comparable numbers are reported here.

Limitations

Parliamentary speech is the whole training distribution: formal, adult, mostly male speakers, read/spoken in a plenary setting. Expect the model to be weaker on conversational speech, children's voices, and text formatted unlike parliamentary transcripts (the tokenizer is case- and punctuation-sensitive).

License

Weights: CC-BY 4.0. Derived from ParCzech4Speech (CC-BY 4.0) — attribution to the ÚFAL, Charles University corpus authors.

Contributors

vvolhejn

4 commits

vvolhejn/pocket-tts-czech

Model

0

stars

4

commits

1

linked in READMEs

Aug 25, 2026

updated

czech
pocket-tts
safetensors
text-to-speech
tts

README

Pocket TTS — Czech (6 layers)

A Czech voice-cloning TTS model for Pocket TTS. Six transformer layers, runs faster than realtime on CPU.

Usage

uvx pocket-tts generate \
  --config hf://vvolhejn/pocket-tts-czech/czech.yaml@7c1fbd0acba765617749dd17f3dbddc2be791cc7 \
  --voice your_voice.wav \
  --text "Dobrý den, toto je český model."

The voice must be an audio file (or a state exported with pocket-tts export-voice from these weights). The named catalog voices — alba, cosette, … — are conditioning states precomputed with the released English weights and will not work here.

Sample voices cut from held-out ParCzech speakers are in voices/:

uvx pocket-tts generate \
  --config hf://vvolhejn/pocket-tts-czech/czech.yaml@7c1fbd0acba765617749dd17f3dbddc2be791cc7 \
  --voice https://huggingface.co/vvolhejn/pocket-tts-czech/resolve/main/voices/cs_m_zenisek.wav \
  --text "Dobrý den, toto je český model."

Training

corpusParCzech4Speech (Czech parliamentary speech, CC-BY 4.0)
training data976 h, 547,597 utterances, 524 speakers
tokenizersentencepiece fitted on the ParCzech transcripts, vocab 3999
teacher24 layers, LSD from scratch, 250k steps, lr 2e-4 constant, flow_batch_multiplier 4
student6 layers, depth-distilled from the teacher's EMA weights, 100k steps, lr 4e-4 cosine, distill_cfg_coef 2.0
weightsEMA (decay 0.9999)
recommended--temperature 0.3, cfg 1 (the default; guidance is baked into the student)

Trained with the Pocket TTS training code, following the non-English recipe.

Evaluation

The Pocket TTS eval pipeline (WER / speaker similarity / UTMOS) is English-only, so no comparable numbers are reported here.

Limitations

Parliamentary speech is the whole training distribution: formal, adult, mostly male speakers, read/spoken in a plenary setting. Expect the model to be weaker on conversational speech, children's voices, and text formatted unlike parliamentary transcripts (the tokenizer is case- and punctuation-sensitive).

License

Weights: CC-BY 4.0. Derived from ParCzech4Speech (CC-BY 4.0) — attribution to the ÚFAL, Charles University corpus authors.

Contributors

vvolhejn

4 commits