CoolFace
Modelpublic

jiangzhuo9357/qwen3-tts-12hz-0-6b-customvoice-gguf

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes82downloads
Model Card

Qwen3-TTS 12Hz 0.6B CustomVoice: synthesize.cpp GGUF

GGUF conversions of the official Qwen3-TTS-12Hz-0.6B-CustomVoice checkpoint for synthesize.cpp.

Ported from QwenLM/Qwen3-TTS revision `022e286b98fbec7e1e916cb940cdf532cd9f488e` and validated on 2026-07-29 against the pinned upstream PyTorch implementation at bfloat16 on CUDA.

A self-contained Qwen3-TTS inference package converted from the official 12 Hz 0.6B CustomVoice checkpoint. It is an autoregressive talker over a residual-vector-quantized codec: a 28-layer Qwen3 decoder emits one semantic code per 80 ms frame, a 5-layer code predictor expands it into fifteen acoustic codes, and a SEANet decoder turns the finished stream into 24 kHz audio. Nine preset Voices ship with it, and the package carries its own byte-pair text frontend, so raw UTF-8 text goes in and audio comes out.

Downloads

ProfileDownloadSizeTensor storageSHA-256What to know before choosing it
BF16qwen3-tts-12hz-0-6b-customvoice-BF16.gguf2274.1 MB (2,274,117,280 bytes)402 BF16 + 255 F3201dfad52dd507c26a14d101c4247d375257aa63b07e62706ec3daa0a33ea515d—
F16qwen3-tts-12hz-0-6b-customvoice-F16.gguf2274.3 MB (2,274,279,584 bytes)266 F16 + 391 F32f79ada9173fbb00b23d2ac2bc859cde75ce4dc135d2710d0f65f00c150113064—
Q8_MIXEDqwen3-tts-12hz-0-6b-customvoice-Q8_MIXED.gguf1425.2 MB (1,425,178,784 bytes)266 Q8_0 + 391 F32d1f9be7b312b42ff90da5146aa0306824ae2870001e05e250461844c3807ca58Only the autoregressive half is quantized. The codec's convolutions carry the kernel in their fastest dimension -- rows of 7, 16, 3 and 1 -- which no block size divides, so it stays at F32. That is 80 per cent of the package quantized for a 37 per cent reduction, and synthesis becomes faster than real time on the machine this was measured on.

All profiles use the same Qwen3-TTS architecture and public synthesize.cpp API. The profile name describes a versioned storage policy, not the language or Execution Backend.

Validation status

validation_level: port_validated

3 graph stages were replayed for 18 cases on DGX Spark CPU, NVIDIA GB10 CUDA 13.3. Duration structure was exact in every case. One of this package's three stages runs on CUDA and two never do. The codec is the half that moves, and it gets same-named accelerator-resident twins the catalog binds against; the talker and the code predictor stay on the CPU on every Execution Backend, because each sampled code is a discrete value that conditions the next step and therefore the sequence length, which is docs/backends.md's discrete-outputs rule. They get no twins at all -- a second copy of 1.8 GB that can never leave the CPU would be read by nothing. Together the two held stages are 59 percent of wall clock under the F16 profile and the codec is the other 41: moving it takes that stage from 8.44 s to 0.24 s and the end-to-end 37-frame case from 20.69 s to 12.50 s, 35 times on the stage and 1.66 end to end. The move is checked rather than trusted -- the replay validator fails if the talker or the predictor places any node off the CPU, or if the codec leaves any node on it -- and six of the eight probes stay bit-identical to the CPU stage, with only audio.pcm moving, in the sixth decimal of its cosine. Suite-wide node totals have not been published for this family, so none is quoted here.

ProfileCPU waveform cosineDGX Spark CUDA waveform cosine
BF160.999427not_measured
F160.9994270.999404
Q8_MIXED0.999427not_measured

Cosine similarity is the measure, not a sample-wise bound, and the codes are replayed rather than reproduced. The oracle draws its codes from PyTorch's generator and this port from its own seeded stream, so reproducing the same codes is neither possible nor the claim: parity replays the oracle's codes and compares everything downstream of the draw on identical inputs.

The deep talker layers carry outlier channels ahead of the final norm — a max-abs of 17.4 at layer 27 while cosine holds at 0.99992 — so a sample-wise threshold would report catastrophic failure on a correct port. The figure quoted is the worst case over the suite. The oracle ran bfloat16 on CUDA and the CPU column runs F32, so these cover a dtype and a device difference as well as an implementation one.

Quality evaluation has not been run. These results establish that the port, Voice selection, deterministic request path, and CPU/CUDA execution work. They do not claim perceptual equivalence, naturalness, intelligibility, or speaker similarity.

A listening audit found no obvious regression. One maintainer compared a small set against the reference and reported nothing audible. That is release evidence, not a measurement: no rated comparison, no panel, no score, and it does not change the validation level. It says a defect large enough to hear was not found in what was heard.

Voices and input

This package exposes 9 preset speaker IDs, preset-catalog. It produces 24000 Hz mono F32 audio. No default speaker is invented; every request must select a Voice.

This package accepts raw UTF-8 text through the built-in synthesize.qwen_bpe frontend, and also accepts exact token IDs already produced by the same vocabulary. The frontend tokenizes text directly: no grapheme-to-phoneme conversion happens or is needed, and the runtime does not silently invoke eSpeak or download a frontend.

Usage

Build synthesize.cpp and synthesize a deterministic request:

bash
git clone https://github.com/handy-computer/synthesize.cpp.git
cd synthesize.cpp
cmake -S . -B build -DSYNTH_BUILD_CLI=ON
cmake --build build -j

hf download jiangzhuo9357/qwen3-tts-12hz-0-6b-customvoice-gguf qwen3-tts-12hz-0-6b-customvoice-F16.gguf \
  --local-dir models/qwen3-tts-12hz-0-6b-customvoice

build/bin/synthesize-cli \
  --model models/qwen3-tts-12hz-0-6b-customvoice/qwen3-tts-12hz-0-6b-customvoice-F16.gguf \
  --output output.wav \
  --text "Qwen3-TTS is awesome!" \
  --language en \
  --voice aiden \
  --seed 0

The same local GGUF can be loaded through the public C ABI and wrapped by C++, Rust, or Python. Model loading never contacts Hugging Face.

License and checkpoint provenance

Both the pinned source at 022e286b and the checkpoint at 85e237c1 carry an explicit Apache-2.0 grant, audited at those revisions rather than at main, with no restriction prose in either card.

Alibaba does not disclose the training corpora, so Apache-2.0 is the basis relied on — the same basis on which Kokoro was accepted.

Both the source and the checkpoint carry an explicit Apache-2.0 grant, so no redistribution assumption is required for this variant.

Only English is declared. The checkpoint also carries codec language tokens for German, Spanish, Chinese, Japanese, French, Korean, Russian, Italian and Brazilian Portuguese, and two of its nine speakers pin a Chinese dialect; those paths load and run, but no language beyond en has its own validation cases, so none is advertised.

The speech tokenizer's encoder half is deliberately not carried. Synthesis runs one way — codes to audio — so nothing in this package can reach it, and its first sixteen codebooks duplicated the decoder's exactly. Dropping it removed 161 tensors and 225 MB.


Original upstream project card

Reproduced from the pinned upstream repository card for offline provenance. The upstream repository remains authoritative.

Qwen3-TTS-12Hz-0.6B-CustomVoice

Qwen3-TTS is a series of advanced multilingual, controllable, robust, and streaming text-to-speech models developed by the Qwen team.

This specific checkpoint is the 0.6B CustomVoice variant, based on the 12Hz tokenizer. It supports 9 premium timbres and allows for fine-grained style control over target voices via natural language instructions across 10 major languages.

Key Features

  • —Multilingual Synthesis: Supports Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian.
  • —Intelligent Control: Adapts tone, rhythm, and emotional expression based on natural language instructions (e.g., "Speak in a very happy tone").
  • —Low Latency: Optimized for streaming generation with the Qwen3-TTS-Tokenizer-12Hz, achieving end-to-end synthesis latency as low as 97ms.

Quickstart

To use Qwen3-TTS, you can install the qwen-tts package:

bash
pip install -U qwen-tts

Sample Usage

python
import torch
import soundfile as sf
from qwen_tts import Qwen3TTSModel

# Load the model
model = Qwen3TTSModel.from_pretrained(
    "Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice",
    device_map="cuda:0",
    dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
)

# Generate speech with specific instructions
wavs, sr = model.generate_custom_voice(
    text="其实我真的有发现,我是一个特别善于观察别人情绪的人。",
    language="Chinese", 
    speaker="Vivian",
    instruct="用特别愤怒的语气说", 
)

# Save the generated audio
sf.write("output_custom_voice.wav", wavs[0], sr)

Supported Speakers

For Qwen3-TTS-12Hz-0.6B-CustomVoice, the following speakers are supported. We recommend using each speaker’s native language for the best results:

SpeakerVoice DescriptionNative Language
VivianBright young female voice.Chinese
SerenaWarm, gentle young female voice.Chinese
Uncle_FuSeasoned male voice, mellow timbre.Chinese
DylanYouthful Beijing male voice.Chinese (Beijing)
EricLively Chengdu male voice.Chinese (Sichuan)
RyanDynamic male voice with rhythm.English
AidenSunny American male voice.English
Ono_AnnaPlayful Japanese female voice.Japanese
SoheeWarm Korean female voice.Korean

Citation

If you find Qwen3-TTS useful for your research, please consider citing:

bibtex
@article{Qwen3-TTS,
  title={Qwen3-TTS Technical Report},
  author={Hangrui Hu and Xinfa Zhu and Ting He and Dake Guo and Bin Zhang and Xiong Wang and Zhifang Guo and Ziyue Jiang and Hongkun Hao and Zishan Guo and Xinyu Zhang and Pei Zhang and Baosong Yang and Jin Xu and Jingren Zhou and Junyang Lin},
  journal={arXiv preprint arXiv:2601.15621},
  year={2026}
}