hugging-apps/ice-012-audio-tts

Space

🧊 ICE-012-Audio β€” multilingual TTS & voice cloning

5

3 commits

1 linked in READMEs

updated Sep 2, 2026

See the code

README

🧊 ICE-012-Audio β€” multilingual TTS & voice cloning

A Gradio demo for darkps/ice-012-audio, a masked-diffusion (non-autoregressive) text-to-speech model with a Qwen3 backbone, an acoustic prosody adapter, and the Higgs-Audio-v2 24 kHz neural codec.

What you can do

  • Synthesize speech in any of the model's ~590 supported languages / dialects.
  • Steer the voice with instruct tags: gender, age, pitch, whisper style, English accent or Chinese regional dialect.
  • Clone a voice from a short reference clip (a transcript is auto-generated with Whisper large-v3-turbo when you don't supply one).
  • Tune speed, diffusion steps and classifier-free guidance under Advanced settings.

Runs on ZeroGPU.

Example audio credits

  • examples/ljspeech_female.wav β€” LJSpeech (public domain).
  • examples/librispeech_male.flac β€” LibriSpeech dev-clean utterance 1272-128104-0003, from LibriSpeech (CC-BY-4.0).

Responsible use

Voice cloning should only be used with audio you own or have explicit permission to process. Do not use this demo to impersonate real people.

gradio
mcp-server

Contributors

multimodalart

3 commits

hugging-apps/ice-012-audio-tts

Space

🧊 ICE-012-Audio β€” multilingual TTS & voice cloning

5

3 commits

1 linked in READMEs

updated Sep 2, 2026

See the code

README

🧊 ICE-012-Audio β€” multilingual TTS & voice cloning

A Gradio demo for darkps/ice-012-audio, a masked-diffusion (non-autoregressive) text-to-speech model with a Qwen3 backbone, an acoustic prosody adapter, and the Higgs-Audio-v2 24 kHz neural codec.

What you can do

  • Synthesize speech in any of the model's ~590 supported languages / dialects.
  • Steer the voice with instruct tags: gender, age, pitch, whisper style, English accent or Chinese regional dialect.
  • Clone a voice from a short reference clip (a transcript is auto-generated with Whisper large-v3-turbo when you don't supply one).
  • Tune speed, diffusion steps and classifier-free guidance under Advanced settings.

Runs on ZeroGPU.

Example audio credits

  • examples/ljspeech_female.wav β€” LJSpeech (public domain).
  • examples/librispeech_male.flac β€” LibriSpeech dev-clean utterance 1272-128104-0003, from LibriSpeech (CC-BY-4.0).

Responsible use

Voice cloning should only be used with audio you own or have explicit permission to process. Do not use this demo to impersonate real people.

gradio
mcp-server

Contributors

multimodalart

3 commits