hugging-apps/ice-012-audio-tts
5
๐ง ICE-012-Audio โ multilingual TTS & voice cloning
A Gradio demo for `darkps/ice-012-audio`, a masked-diffusion (non-autoregressive) text-to-speech model with a Qwen3 backbone, an acoustic prosody adapter, and the Higgs-Audio-v2 24 kHz neural codec.
What you can do
- Synthesize speech in any of the model's ~590 supported languages / dialects.
- Steer the voice with instruct tags: gender, age, pitch, whisper style, English accent or Chinese regional dialect.
- Clone a voice from a short reference clip (a transcript is auto-generated with Whisper large-v3-turbo when you don't supply one).
- Tune speed, diffusion steps and classifier-free guidance under Advanced settings.
Runs on ZeroGPU.
Example audio credits
examples/ljspeech_female.wavโ LJSpeech (public domain).examples/librispeech_male.flacโ LibriSpeechdev-cleanutterance1272-128104-0003, from LibriSpeech (CC-BY-4.0).
Responsible use
Voice cloning should only be used with audio you own or have explicit permission to process. Do not use this demo to impersonate real people.
