cstr/qwen3-tts-0.6b-customvoice-GGUF
Qwen3-TTS-12Hz-0.6B-CustomVoice — GGUF (ggml-quantised)
GGUF / ggml conversion of `Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice` for use with [CrispStrobe/CrispASR](https://github.com/CrispStrobe/CrispASR).
The CustomVoice variant is the fixed-speaker sibling of Qwen3-TTS-12Hz-0.6B-Base: instead of cloning a voice from a 3 s reference WAV (the Base path), it ships 9 baked speaker tokens picked via a --voice <name> flag. No ECAPA forward, no codec encoder, no reference audio required. Two of the speakers (dylan, eric) carry Chinese-dialect overrides (Beijing / Sichuan) that re-route language_id when synthesising Chinese-or-auto.
Pair this with the codec at `cstr/qwen3-tts-tokenizer-12hz-GGUF` — the talker emits 16-codebook RVQ codes that the codec decoder renders to 24 kHz PCM.
Files
The 1.7B-CustomVoice variant ships at `cstr/qwen3-tts-1.7b-customvoice-GGUF`. For ICL voice cloning use `cstr/qwen3-tts-0.6b-base-GGUF` instead.
Quick start
# 1. Build CrispASR
git clone https://github.com/CrispStrobe/CrispASR
cd CrispASR
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j --target crispasr
# 2. Pull the talker + the codec
huggingface-cli download cstr/qwen3-tts-0.6b-customvoice-GGUF qwen3-tts-12hz-0.6b-customvoice-q8_0.gguf --local-dir .
huggingface-cli download cstr/qwen3-tts-tokenizer-12hz-GGUF qwen3-tts-tokenizer-12hz.gguf --local-dir .
# 3. Synthesise — pick a speaker by name
./build/bin/crispasr --backend qwen3-tts-customvoice \
-m qwen3-tts-12hz-0.6b-customvoice-q8_0.gguf \
--codec-model qwen3-tts-tokenizer-12hz.gguf \
--voice ryan \
--tts "Hello, this is the Ryan speaker." \
--tts-output ryan.wavFor auto-download simply pass -m auto:
./build/bin/crispasr --backend qwen3-tts-customvoice -m auto \
--voice serena \
--tts "Auto-download fetches both files." \
--tts-output out.wavQuality verification
ASR roundtrip via `cstr/parakeet-tdt-0.6b-v3-GGUF`:
Architecture
CustomVoice differs from Base in that the speaker embedding is not computed via ECAPA on a reference WAV. Instead talker.get_input_embeddings()(spk_id) retrieves a row from the codec embedding table directly (modeling_qwen3_tts.py:2091). No reference audio, no codec encoder needed at runtime.
CrispASR backend integration
The runtime branches on qwen3tts.tts_model_type ("customvoice" vs "base" vs "voicedesign") so the same backend object handles all three Qwen3-TTS variants.
Conversion
python models/convert-qwen3-tts-to-gguf.py \
--input Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice \
--output qwen3-tts-12hz-0.6b-customvoice-f16.gguf \
--outtype f16
build/bin/crispasr-quantize qwen3-tts-12hz-0.6b-customvoice-f16.gguf \
qwen3-tts-12hz-0.6b-customvoice-q8_0.gguf q8_0The converter writes qwen3tts.tts_model_type=custom_voice plus qwen3tts.spk_names, qwen3tts.spk_token_ids, and qwen3tts.spk_dialect_token_ids so the runtime can resolve --voice <name> to the right speaker token + dialect override at load time.
Attribution
- Talker base: `Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice` (Apache 2.0).
- Codec base: `Qwen/Qwen3-TTS-Tokenizer-12Hz` (Apache 2.0) — see `cstr/qwen3-tts-tokenizer-12hz-GGUF`.
- Reference engine: QwenLM/Qwen3-TTS (
modeling_qwen3_tts.py). - GGUF conversion + ggml runtime: `CrispStrobe/CrispASR` — see
src/qwen3_tts.cpp,models/convert-qwen3-tts-to-gguf.py.
License
Apache 2.0 — both the upstream talker and codec, and the GGUF conversion.
Provenance and EU AI Act Art. 53 note
- Upstream model: Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice — published by
Qwen. - Upstream licence:
apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not. - What was done here: format conversion and/or quantisation only (GGUF/GGML). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
- Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository. No training-content summary was found on the upstream model card at the time of writing; that documentation gap is upstream's and is not filled here.
- Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.
