digibuzz24/Qwen3-TTS-GGUF-fork
Qwen3-TTS GGUF
GGUF weights for qwentts.cpp, a C++17/GGML port of Qwen3-TTS 12 Hz (Qwen team, Alibaba). Multilingual zero shot TTS with named speakers and Mandarin dialects, 24 kHz mono. Runs on CPU, CUDA, Metal, Vulkan.
Files
Two GGUFs load together :
qwen-talker-{size}-{mode}-{variant}.gguf Qwen3 LM + code predictor MTP head + optional speaker encoder, text -> 12 Hz codes qwen-tokenizer-12hz-{variant}.gguf SEANet + ConvNeXt + DAC v2 + RVQ, 12 Hz codes <-> 24 kHz audio
Three modes are available across two talker sizes :
The tokenizer is shared across every talker.
Quick start
git clone --recurse-submodules https://github.com/ServeurpersoCom/qwentts.cpp.git
cd qwentts.cpp && ./buildcuda.sh
mkdir -p models
huggingface-cli download Serveurperso/Qwen3-TTS-GGUF \
qwen-talker-1.7b-base-Q8_0.gguf qwen-tokenizer-12hz-Q8_0.gguf \
--local-dir models
cd examples
./base.sh # named speaker -> base.wav
./clone.sh # voice cloning -> clone.wav
./customvoice.sh # custom voice mode -> customvoice.wav
./tts.sh # voice design -> tts.wavBackends
Set GGML_BACKEND to force a device, otherwise the runtime picks the best one available.
Quantization policy
Tokenizer GGUFs are not uniform quants. Three categories get a dedicated treatment :
Conv kernel rows (K=7,3,1) never divide a K-quant block size, so the quantizer skips the Q* intermediates and lands on F16 directly. This is the last resort branch of llama.cpp's tensor_type_fallback applied unconditionally for these kernels. F16 has no block size and matches the runtime target dtype on every backend. The talker LM (Qwen3 backbone, hidden divisible by 256) follows standard llama.cpp K-quant across variants. The code predictor MTP head and the speaker encoder live in the talker GGUF and share its quantization.
License
Upstream model : Qwen3-TTS by Alibaba / Qwen team, Apache 2.0 Audio codec : Qwen3-TTS-Tokenizer-12Hz (Qwen team), Apache 2.0 GGUF tooling : qwentts.cpp, MIT
