CoolFace
Modelpublic

cstr/qwen3-tts-0.6b-customvoice-GGUF

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
4likes2.6kdownloads
Model Card

Qwen3-TTS-12Hz-0.6B-CustomVoice — GGUF (ggml-quantised)

GGUF / ggml conversion of `Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice` for use with [CrispStrobe/CrispASR](https://github.com/CrispStrobe/CrispASR).

The CustomVoice variant is the fixed-speaker sibling of Qwen3-TTS-12Hz-0.6B-Base: instead of cloning a voice from a 3 s reference WAV (the Base path), it ships 9 baked speaker tokens picked via a --voice <name> flag. No ECAPA forward, no codec encoder, no reference audio required. Two of the speakers (dylan, eric) carry Chinese-dialect overrides (Beijing / Sichuan) that re-route language_id when synthesising Chinese-or-auto.

SpeakerLanguage / dialect
aiden (default)English (M)
dylanBeijing dialect (M, dialect_token=2074)
ericSichuan dialect (M, dialect_token=2062)
ono_annaEnglish (F)
ryanEnglish (M)
serenaEnglish (F)
soheeEnglish (F)
uncle_fuEnglish (M, older)
vivianEnglish (F)

Pair this with the codec at `cstr/qwen3-tts-tokenizer-12hz-GGUF` — the talker emits 16-codebook RVQ codes that the codec decoder renders to 24 kHz PCM.

Files

FileQuantSizeNotes
qwen3-tts-12hz-0.6b-customvoice-q8_0.ggufQ8_0968 MBRecommended — ASR-roundtrip word-exact vs reference

The 1.7B-CustomVoice variant ships at `cstr/qwen3-tts-1.7b-customvoice-GGUF`. For ICL voice cloning use `cstr/qwen3-tts-0.6b-base-GGUF` instead.

Quick start

bash
# 1. Build CrispASR
git clone https://github.com/CrispStrobe/CrispASR
cd CrispASR
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j --target crispasr

# 2. Pull the talker + the codec
huggingface-cli download cstr/qwen3-tts-0.6b-customvoice-GGUF qwen3-tts-12hz-0.6b-customvoice-q8_0.gguf --local-dir .
huggingface-cli download cstr/qwen3-tts-tokenizer-12hz-GGUF qwen3-tts-tokenizer-12hz.gguf --local-dir .

# 3. Synthesise — pick a speaker by name
./build/bin/crispasr --backend qwen3-tts-customvoice \
    -m qwen3-tts-12hz-0.6b-customvoice-q8_0.gguf \
    --codec-model qwen3-tts-tokenizer-12hz.gguf \
    --voice ryan \
    --tts "Hello, this is the Ryan speaker." \
    --tts-output ryan.wav

For auto-download simply pass -m auto:

bash
./build/bin/crispasr --backend qwen3-tts-customvoice -m auto \
    --voice serena \
    --tts "Auto-download fetches both files." \
    --tts-output out.wav

Quality verification

ASR roundtrip via `cstr/parakeet-tdt-0.6b-v3-GGUF`:

SpeakerSynthesised textParakeet output
vivian"Hello, this is a CustomVoice test using the vivian speaker.""Hello! This is a custom voice test using the Vivian speaker." (verbatim modulo case/punct)
aiden"The quick brown fox jumps over the lazy dog.""The quick brown fox jumps over the lazy dog." (verbatim)
serena"Testing the new backend alias and the serena speaker.""Testing the new back end Ilias and the Serena speaker." (1 ASR misrecognition; audio clean)
dylan"你好,今天天气真不错。"dialect override engaged (language_id=2074); 3.28 s clean audio

Architecture

ComponentDetails
Talker LMQwen3 (28 layers, 1024 hidden, 16 heads, 8 KV heads, head_dim=64)
Output head16 codebooks × 2048 (RVQ) — emits codes for the codec
Code predictor5L Qwen3 + 15 separate codecembedding/lmhead pairs (top-k=50, temp=0.9 sampling)
CodecQwen3-TTS-Tokenizer-12Hz (separate GGUF, 12.5 fps RVQ)
Audio24 kHz mono float32 PCM

CustomVoice differs from Base in that the speaker embedding is not computed via ECAPA on a reference WAV. Instead talker.get_input_embeddings()(spk_id) retrieves a row from the codec embedding table directly (modeling_qwen3_tts.py:2091). No reference audio, no codec encoder needed at runtime.

CrispASR backend integration

FlagPurpose
--backend qwen3-tts-customvoiceSelects the CustomVoice runtime path
--voice <name>Picks a fixed speaker by name (default: first speaker, aiden)
--tts "..."Text to synthesise
--codec-model PATHQwen3-TTS-Tokenizer-12Hz GGUF
--tts-output PATHOutput WAV (24 kHz mono)

The runtime branches on qwen3tts.tts_model_type ("customvoice" vs "base" vs "voicedesign") so the same backend object handles all three Qwen3-TTS variants.

Conversion

bash
python models/convert-qwen3-tts-to-gguf.py \
    --input Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice \
    --output qwen3-tts-12hz-0.6b-customvoice-f16.gguf \
    --outtype f16

build/bin/crispasr-quantize qwen3-tts-12hz-0.6b-customvoice-f16.gguf \
    qwen3-tts-12hz-0.6b-customvoice-q8_0.gguf q8_0

The converter writes qwen3tts.tts_model_type=custom_voice plus qwen3tts.spk_names, qwen3tts.spk_token_ids, and qwen3tts.spk_dialect_token_ids so the runtime can resolve --voice <name> to the right speaker token + dialect override at load time.

Attribution

License

Apache 2.0 — both the upstream talker and codec, and the GGUF conversion.

Provenance and EU AI Act Art. 53 note

  • —Upstream model: Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice — published by Qwen.
  • —Upstream licence: apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not.
  • —What was done here: format conversion and/or quantisation only (GGUF/GGML). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
  • —Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository. No training-content summary was found on the upstream model card at the time of writing; that documentation gap is upstream's and is not filled here.
  • —Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.