CoolFace
Modelpublic

jmtl/hk-cantonese-cc0-qwen3-tts-GGUF

sourceHugging Faceapache-2.0updated 20d agoView on Hugging Face
0likes36downloads
Model Card

hk-cantonese-cc0 — GGUF (v0.0.1)

q8_0 GGUF of hk-cantonese-cc0, a Hong Kong Cantonese female voice for Qwen3-TTS 12Hz 1.7B. Runs on qwen3-tts.cpp — no Python, no PyTorch. The adapter form, the evaluation and the training recipe are in the main repository; see Which model this is a conversion of below, because it is not a conversion of that repository's adapter.

Status: early usable release (v0.0.1). Single donor voice.

Files

qwen-talker-hk-cantonese-cc0-v0.0.1-q8_0.gguf   2.0 GB   The voice
qwen-tokenizer-12hz-q8_0.gguf                   278 MB   12 Hz codec, for --codec-model

Verify your download:

4278d9944b06013d6f8cb33358e2eacb  qwen-talker-hk-cantonese-cc0-v0.0.1-q8_0.gguf
39819bc6f94b35837b94b22cdd0ce95e  qwen-tokenizer-12hz-q8_0.gguf

The codec is an unmodified conversion of Qwen/Qwen3-TTS-Tokenizer-12Hz (Apache-2.0), shipped so this repository is self-contained. Any other build of the same tokenizer works in its place.

Which model this is a conversion of

The original merged release, not a rebuild of the sibling adapter. The two are different files:

4278d9944b06013d6f8cb33358e2eacb   from the merged release   <- published here
8a89a053d0774f0b8248bf17fc4aa45a   from an adapter rebuild   <- not published

If you rebuild with apply_adapter.py and convert it yourself, expect the second digest. They were compared blind on a 629-character read as the two arms of the same listening test: the adapter rebuild marked 8 characters against the merged release's 11, verdict not clearly worse. Equivalent by ear, measurably different as files.

Build the runtime

bash
git clone https://github.com/Danmoreng/qwen3-tts.cpp.git
cd qwen3-tts.cpp
git checkout 16bb5afcd06311031c72a8488f8d59660dc2fb46
cmake -B build -DGGML_CUDA=ON      # omit -DGGML_CUDA=ON for a CPU-only build
cmake --build build -j

This is the commit these weights were rendered and evaluated on. Build on a native filesystem — under WSL a CUDA build over /mnt is roughly 138x slower.

Run

-m takes the directory holding the talker; --codec-model takes the codec file.

bash
mkdir -p talker && mv qwen-talker-hk-cantonese-cc0-v0.0.1-q8_0.gguf talker/

./build/qwen3-tts-cli \
  -m ./talker \
  --codec-model ./qwen-tokenizer-12hz-q8_0.gguf \
  --speaker hk_cantonese_speaker \
  --seed 20260905 \
  -t "今日天氣好好,我哋出去行下街啦。" \
  -o out.wav
  • —--speaker hk_cantonese_speaker selects the baked voice. This build conditions on a baked speaker rather than reference audio, so it cannot voice-clone.
  • —--seed makes a render repeatable. Without it the same text gives a different draw.
  • —The runtime finds the talker by matching qwen-talker in the filename, so any name keeping that prefix works.

A single render spends about 6 s of wall clock on roughly 1.2 s of synthesis; the rest is process start and model load. There is no batch mode in this build: the CLI takes exactly one -t per process, and --bench-server re-synthesises a single string rather than reading a list, so a passage pays that load per clip.

The CLI writes WAV. Convert to FLAC if the output will be kept or compared — treat the WAV as scratch.

Operational constraints

  • —Chunk to ~23 characters, split on ,。:;!?. Rejection rates rise from 17% at ≤ 26 characters to 43% above 41.
  • —Keep default sampling. --temperature 0, --temperature 0.6, --top-p 0.90 and --top-p 0.85 all produce runaway loops.
  • —Trim the edges. Of 27 clips measured, 24 carried over a second of silence at an edge — lead median 1.08 s out to 9.33 s, tail median 1.34 s out to 5.40 s. Trim before using clip duration as a metric, and keep the raw duration if you do.
  • —Detect stalls. About 5% of renders come back as silence; discard anything whose non-silent ratio is ≤ 0.14 and re-roll.
  • —Colloquial Cantonese only. Standard written Chinese (書面語) and literary readings mispronounce.

Failure modes, evaluation and the training recipe are in the main repository.

Licence

Apache-2.0, matching Qwen3-TTS 12Hz 1.7B, which these weights derive from. The voice is one pseudonymous Common Voice contributor's, dedicated under CC-0.

Do not label this voice as synthetic or anonymous, do not attempt real-world re-identification, and do not use it for impersonation.