jmtl/hk-cantonese-cc0-qwen3-tts-GGUF
hk-cantonese-cc0 — GGUF (v0.0.1)
q8_0 GGUF of hk-cantonese-cc0, a Hong Kong Cantonese female voice for Qwen3-TTS 12Hz 1.7B. Runs on qwen3-tts.cpp — no Python, no PyTorch. The adapter form, the evaluation and the training recipe are in the main repository; see Which model this is a conversion of below, because it is not a conversion of that repository's adapter.
Status: early usable release (v0.0.1). Single donor voice.
Files
qwen-talker-hk-cantonese-cc0-v0.0.1-q8_0.gguf 2.0 GB The voice
qwen-tokenizer-12hz-q8_0.gguf 278 MB 12 Hz codec, for --codec-modelVerify your download:
4278d9944b06013d6f8cb33358e2eacb qwen-talker-hk-cantonese-cc0-v0.0.1-q8_0.gguf
39819bc6f94b35837b94b22cdd0ce95e qwen-tokenizer-12hz-q8_0.ggufThe codec is an unmodified conversion of Qwen/Qwen3-TTS-Tokenizer-12Hz (Apache-2.0), shipped so this repository is self-contained. Any other build of the same tokenizer works in its place.
Which model this is a conversion of
The original merged release, not a rebuild of the sibling adapter. The two are different files:
4278d9944b06013d6f8cb33358e2eacb from the merged release <- published here
8a89a053d0774f0b8248bf17fc4aa45a from an adapter rebuild <- not publishedIf you rebuild with apply_adapter.py and convert it yourself, expect the second digest. They were compared blind on a 629-character read as the two arms of the same listening test: the adapter rebuild marked 8 characters against the merged release's 11, verdict not clearly worse. Equivalent by ear, measurably different as files.
Build the runtime
git clone https://github.com/Danmoreng/qwen3-tts.cpp.git
cd qwen3-tts.cpp
git checkout 16bb5afcd06311031c72a8488f8d59660dc2fb46
cmake -B build -DGGML_CUDA=ON # omit -DGGML_CUDA=ON for a CPU-only build
cmake --build build -jThis is the commit these weights were rendered and evaluated on. Build on a native filesystem — under WSL a CUDA build over /mnt is roughly 138x slower.
Run
-m takes the directory holding the talker; --codec-model takes the codec file.
mkdir -p talker && mv qwen-talker-hk-cantonese-cc0-v0.0.1-q8_0.gguf talker/
./build/qwen3-tts-cli \
-m ./talker \
--codec-model ./qwen-tokenizer-12hz-q8_0.gguf \
--speaker hk_cantonese_speaker \
--seed 20260905 \
-t "今日天氣好好,我哋出去行下街啦。" \
-o out.wav--speaker hk_cantonese_speakerselects the baked voice. This build conditions on a baked speaker rather than reference audio, so it cannot voice-clone.--seedmakes a render repeatable. Without it the same text gives a different draw.- The runtime finds the talker by matching
qwen-talkerin the filename, so any name keeping that prefix works.
A single render spends about 6 s of wall clock on roughly 1.2 s of synthesis; the rest is process start and model load. There is no batch mode in this build: the CLI takes exactly one -t per process, and --bench-server re-synthesises a single string rather than reading a list, so a passage pays that load per clip.
The CLI writes WAV. Convert to FLAC if the output will be kept or compared — treat the WAV as scratch.
Operational constraints
- Chunk to ~23 characters, split on ,。:;!?. Rejection rates rise from 17% at ≤ 26 characters to 43% above 41.
- Keep default sampling.
--temperature 0,--temperature 0.6,--top-p 0.90and--top-p 0.85all produce runaway loops. - Trim the edges. Of 27 clips measured, 24 carried over a second of silence at an edge — lead median 1.08 s out to 9.33 s, tail median 1.34 s out to 5.40 s. Trim before using clip duration as a metric, and keep the raw duration if you do.
- Detect stalls. About 5% of renders come back as silence; discard anything whose non-silent ratio is ≤ 0.14 and re-roll.
- Colloquial Cantonese only. Standard written Chinese (書面語) and literary readings mispronounce.
Failure modes, evaluation and the training recipe are in the main repository.
Licence
Apache-2.0, matching Qwen3-TTS 12Hz 1.7B, which these weights derive from. The voice is one pseudonymous Common Voice contributor's, dedicated under CC-0.
Do not label this voice as synthetic or anonymous, do not attempt real-world re-identification, and do not use it for impersonation.
