jmtl/hk-cantonese-cc0-qwen3-tts
hk-cantonese-cc0 (v0.0.1)
A 140 MB LoRA adapter for Qwen3-TTS 12Hz 1.7B that provides a Hong Kong Cantonese female voice. Trained exclusively on CC-0 data and distributed under Apache-2.0.
Status: early usable release (v0.0.1). Single donor voice. For comprehensive benchmarks and training logs, see `docs/MODEL_CARD.md`.
Files
adapter_model.safetensors 139.5 MB Rank-64 LoRA over 196 talker projections (fp16)
extra_tensors.safetensors 8 KB Baked speaker embedding, codec index 3000
adapter_config.json LoRA config (r = 64, alpha = 64)
SHA256SUMS Digests of the two .safetensors files
LICENSE, NOTICE Apache-2.0 and required attribution
scripts/
apply_adapter.py Merge script; builds the runnable model
docs/
MODEL_CARD.md Failure modes, empirical bounds, evaluation
recovery_report.json Per-tensor SVD reconstruction error reportThe three adapter files stay at the repository root, where a loader expects them.
Quickstart
1. Build the weights
Download Qwen/Qwen3-TTS-12Hz-1.7B-Base (requires the -Base variant, not -CustomVoice or -VoiceDesign).
# Standard inference build (~35s on GPU, ~8 GB system RAM on CPU)
python scripts/apply_adapter.py --base ./Qwen3-TTS-12Hz-1.7B-Base --adapter . --out ./hk-cantonese
# Add --trainable ONLY if you intend to run further fine-tuning
python scripts/apply_adapter.py --base ./Qwen3-TTS-12Hz-1.7B-Base --adapter . --out ./hk-cantonese-trainable --trainableThe two builds are not interchangeable. The inference build sets tts_model_type to custom_voice and drops the 76 speaker-encoder tensors, so training from it fails at the first batch with TypeError: 'NoneType' object is not callable.
Note: standardpeftloaders cannot load this directly, because the voice relies on a non-LoRA baked embedding row (extra_tensors.safetensors) and the build drops the upstream speaker encoder.scripts/apply_adapter.pyneeds onlytorchandsafetensors. Run it from the repository root so--adapter .resolves.
2. Inference
import json, torch, soundfile as sf
from qwen_tts import Qwen3TTSModel
path = "./hk-cantonese"
m = Qwen3TTSModel.from_pretrained(
path,
dtype=torch.bfloat16,
device_map="cuda",
attn_implementation="sdpa"
)
# Load baked speaker ID
with open(f"{path}/config.json", encoding="utf-8") as f:
cfg = json.load(f)
speaker = next(iter(cfg["talker_config"]["spk_id"]))
# Split text into short clauses (~23 characters max)
text = "今日天氣好好,我哋出去行下街啦。"
# Keep language="chinese"; the adapter handles Cantonese phonology
wavs, sr = m.generate_custom_voice(
text=[text],
speaker=[speaker],
language=["chinese"]
)
sf.write("out.flac", wavs[0], sr)Operational constraints
- Chunking is mandatory. Keep segments around 23 characters and split on punctuation (,。:;!?). Rejection rates rise from 17% at ≤ 26 characters to 43% above 41, measured over 156 blind trials (Cochran–Armitage p = 0.0032).
- Keep default sampling. T = 0, T = 0.6, top-p 0.90 and top-p 0.85 all produce repetitive runaway loops. Values between those and the defaults are untested, so the defaults are the only known-good configuration.
- Trim the edges. Of 27 clips measured, 24 carried over a second of silence at an edge — lead median 1.08 s out to 9.33 s, tail median 1.34 s out to 5.40 s. The snippet above writes the raw clip, so trim before using duration as a metric and keep the raw duration beside the trimmed one.
- Detect stalls. Approximately 5% of runs emit silence. Discard outputs whose non-silent audio ratio is ≤ 0.14; valid speech runs at ≥ 0.32, with no overlap across 94 calibration clips.
- Colloquial Cantonese only. Trained on spoken Hong Kong Cantonese (口語). Standard written Chinese (書面語) and formal literary readings produce mispronunciations. On text in its own register, a native listener marked 9 characters wrong in a 1,377-character read (0.65%).
- Not a voice-cloning base. New target speakers fine-tuned on top of this checkpoint bleed into this donor's timbre; ECAPA recovers the inherited donor at 0.521 against a 0.469 threshold.
Training recipe
If rebuilding from scratch:
- Base:
Qwen/Qwen3-TTS-12Hz-1.7B-Base - Language phase: Common Voice 22 (
yue), 33,104 clips (47.3 h). LoRA r = 32, alpha = 64, LR 1e-4, 1 epoch,--keep_speaker_encoder. - Speaker phase: 343 clips from a single HK female contributor (
yue), chained from the language checkpoint. LoRA r = 32, alpha = 64, LR 1e-4, 1 epoch. - Hyperparameters: batch size 1, gradient accumulation 4, warmup 0.03,
sub_talker_weight 0. About 2 h 25 m on an RTX 3080 Ti.
The two chained r = 32 phases are why the recovered adapter is rank 64.
Licence and attribution
Adapter and scripts: Apache-2.0.
Voice donor: dedicated under CC-0 via Mozilla Common Voice by a pseudonymous contributor.
Reproducing the corpus. Name the release: the language phase is Common Voice 22, and the 343-clip speaker set came from an earlier yue release whose number was not recorded. A different release renumbers and re-partitions clips, so selecting the same contributor from another one yields different recordings. Only yue was used, not zh-HK. The linked dataset is a community mirror of release 22 whose validated split matches this pool at 191,152 clips; official releases come from Mozilla directly, as mozilla-foundation on the Hub stops at release 17.
Usage restrictions. Do not label this voice as synthetic or anonymous, do not attempt real-world re-identification, and do not use it for impersonation.
