CoolFace
Apppublic

hugging-apps/ice-012-audio-tts

sourceHugging Faceupdated 23d agoView on Hugging Face
5likes
App README

๐ŸงŠ ICE-012-Audio โ€” multilingual TTS & voice cloning

A Gradio demo for `darkps/ice-012-audio`, a masked-diffusion (non-autoregressive) text-to-speech model with a Qwen3 backbone, an acoustic prosody adapter, and the Higgs-Audio-v2 24 kHz neural codec.

What you can do

  • โ€”Synthesize speech in any of the model's ~590 supported languages / dialects.
  • โ€”Steer the voice with instruct tags: gender, age, pitch, whisper style, English accent or Chinese regional dialect.
  • โ€”Clone a voice from a short reference clip (a transcript is auto-generated with Whisper large-v3-turbo when you don't supply one).
  • โ€”Tune speed, diffusion steps and classifier-free guidance under Advanced settings.

Runs on ZeroGPU.

Example audio credits

  • โ€”examples/ljspeech_female.wav โ€” LJSpeech (public domain).
  • โ€”examples/librispeech_male.flac โ€” LibriSpeech dev-clean utterance 1272-128104-0003, from LibriSpeech (CC-BY-4.0).

Responsible use

Voice cloning should only be used with audio you own or have explicit permission to process. Do not use this demo to impersonate real people.