CoolFace
Modelpublic

jmtl/hk-cantonese-cc0-qwen3-tts

sourceHugging Faceapache-2.0updated 20d agoView on Hugging Face
1likes
Model Card

hk-cantonese-cc0 (v0.0.1)

A 140 MB LoRA adapter for Qwen3-TTS 12Hz 1.7B that provides a Hong Kong Cantonese female voice. Trained exclusively on CC-0 data and distributed under Apache-2.0.

Status: early usable release (v0.0.1). Single donor voice. For comprehensive benchmarks and training logs, see `docs/MODEL_CARD.md`.

Files

adapter_model.safetensors     139.5 MB   Rank-64 LoRA over 196 talker projections (fp16)
extra_tensors.safetensors         8 KB   Baked speaker embedding, codec index 3000
adapter_config.json                      LoRA config (r = 64, alpha = 64)
SHA256SUMS                               Digests of the two .safetensors files
LICENSE, NOTICE                          Apache-2.0 and required attribution
scripts/
  apply_adapter.py                       Merge script; builds the runnable model
docs/
  MODEL_CARD.md                          Failure modes, empirical bounds, evaluation
  recovery_report.json                   Per-tensor SVD reconstruction error report

The three adapter files stay at the repository root, where a loader expects them.

Quickstart

1. Build the weights

Download Qwen/Qwen3-TTS-12Hz-1.7B-Base (requires the -Base variant, not -CustomVoice or -VoiceDesign).

bash
# Standard inference build (~35s on GPU, ~8 GB system RAM on CPU)
python scripts/apply_adapter.py --base ./Qwen3-TTS-12Hz-1.7B-Base --adapter . --out ./hk-cantonese

# Add --trainable ONLY if you intend to run further fine-tuning
python scripts/apply_adapter.py --base ./Qwen3-TTS-12Hz-1.7B-Base --adapter . --out ./hk-cantonese-trainable --trainable

The two builds are not interchangeable. The inference build sets tts_model_type to custom_voice and drops the 76 speaker-encoder tensors, so training from it fails at the first batch with TypeError: 'NoneType' object is not callable.

Note: standard peft loaders cannot load this directly, because the voice relies on a non-LoRA baked embedding row (extra_tensors.safetensors) and the build drops the upstream speaker encoder. scripts/apply_adapter.py needs only torch and safetensors. Run it from the repository root so --adapter . resolves.

2. Inference

python
import json, torch, soundfile as sf
from qwen_tts import Qwen3TTSModel

path = "./hk-cantonese"
m = Qwen3TTSModel.from_pretrained(
    path,
    dtype=torch.bfloat16,
    device_map="cuda",
    attn_implementation="sdpa"
)

# Load baked speaker ID
with open(f"{path}/config.json", encoding="utf-8") as f:
    cfg = json.load(f)
speaker = next(iter(cfg["talker_config"]["spk_id"]))

# Split text into short clauses (~23 characters max)
text = "今日天氣好好,我哋出去行下街啦。"

# Keep language="chinese"; the adapter handles Cantonese phonology
wavs, sr = m.generate_custom_voice(
    text=[text],
    speaker=[speaker],
    language=["chinese"]
)
sf.write("out.flac", wavs[0], sr)

Operational constraints

  • —Chunking is mandatory. Keep segments around 23 characters and split on punctuation (,。:;!?). Rejection rates rise from 17% at ≤ 26 characters to 43% above 41, measured over 156 blind trials (Cochran–Armitage p = 0.0032).
  • —Keep default sampling. T = 0, T = 0.6, top-p 0.90 and top-p 0.85 all produce repetitive runaway loops. Values between those and the defaults are untested, so the defaults are the only known-good configuration.
  • —Trim the edges. Of 27 clips measured, 24 carried over a second of silence at an edge — lead median 1.08 s out to 9.33 s, tail median 1.34 s out to 5.40 s. The snippet above writes the raw clip, so trim before using duration as a metric and keep the raw duration beside the trimmed one.
  • —Detect stalls. Approximately 5% of runs emit silence. Discard outputs whose non-silent audio ratio is ≤ 0.14; valid speech runs at ≥ 0.32, with no overlap across 94 calibration clips.
  • —Colloquial Cantonese only. Trained on spoken Hong Kong Cantonese (口語). Standard written Chinese (書面語) and formal literary readings produce mispronunciations. On text in its own register, a native listener marked 9 characters wrong in a 1,377-character read (0.65%).
  • —Not a voice-cloning base. New target speakers fine-tuned on top of this checkpoint bleed into this donor's timbre; ECAPA recovers the inherited donor at 0.521 against a 0.469 threshold.

Training recipe

If rebuilding from scratch:

  • —Base: Qwen/Qwen3-TTS-12Hz-1.7B-Base
  • —Language phase: Common Voice 22 (yue), 33,104 clips (47.3 h). LoRA r = 32, alpha = 64, LR 1e-4, 1 epoch, --keep_speaker_encoder.
  • —Speaker phase: 343 clips from a single HK female contributor (yue), chained from the language checkpoint. LoRA r = 32, alpha = 64, LR 1e-4, 1 epoch.
  • —Hyperparameters: batch size 1, gradient accumulation 4, warmup 0.03, sub_talker_weight 0. About 2 h 25 m on an RTX 3080 Ti.

The two chained r = 32 phases are why the recovered adapter is rank 64.

Licence and attribution

ComponentSourceLicence
Base modelQwen/Qwen3-TTS-12Hz-1.7B-BaseApache-2.0
Language phaseCommon Voice yue, release 22, 33,104 clips / 47.3 hCC-0
Speaker phaseOne Common Voice yue HK contributor, 343 clipsCC-0

Adapter and scripts: Apache-2.0.

Voice donor: dedicated under CC-0 via Mozilla Common Voice by a pseudonymous contributor.

Reproducing the corpus. Name the release: the language phase is Common Voice 22, and the 343-clip speaker set came from an earlier yue release whose number was not recorded. A different release renumbers and re-partitions clips, so selecting the same contributor from another one yields different recordings. Only yue was used, not zh-HK. The linked dataset is a community mirror of release 22 whose validated split matches this pool at 191,152 clips; official releases come from Mozilla directly, as mozilla-foundation on the Hub stops at release 17.

Usage restrictions. Do not label this voice as synthetic or anonymous, do not attempt real-world re-identification, and do not use it for impersonation.