RumiLabs/Qwen3-TTS-12Hz-1.7B-CustomVoice-MLX-4bit
Qwen3-TTS-12Hz-1.7B-CustomVoice — MLX INT4 LLM + FP16 vocoder
Private RumiLabs MLX build of Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice for on-device Apple Silicon inference. This is the shipping target for the Yobi/rumi TTS pipeline, superseding `RumiLabs/Qwen3-TTS-12Hz-1.7B-VoiceDesign-MLX-4bit`.
Why CustomVoice, not VoiceDesign
VoiceDesign treats instruct as a whole-voice description (identity + emotion mixed). Two different emotion prompts produce two different speakers — which we verified empirically: 5 emotion templates → 4 different female voices + 1 male, none anchored to a reference. The --voice parameter is a silent no-op on VoiceDesign and --ref_audio cloning is too weak to anchor identity.
CustomVoice decouples them: voice = preset speaker (9 timbres), instruct = emotion/style only. Listening A/B confirmed: same speaker (Aiden) across all 5 emotions, quality matches VoiceDesign, emotions clearly distinct. One-line code change, no training.
See qwen3_tts.py:1071-1078 for the variant API differences.
What's in this bundle
- LLM:
int4 group-64 affineviamlx-audioconvert. - `speech_tokenizer/` (RVQ codec + Code2Wav): cast from FP32 → FP16 (682 MB → 341 MB). FP16 retains full audible fidelity (40–61 dB SNR on 5-emotion A/B).
- Bundle on disk: 1.80 GiB
- Peak resident on M3 Ultra: 5.24 GiB (+240 MB vs VoiceDesign — CustomVoice's preset-speaker tables add weight).
- 9 preset speakers included (English: Ryan, Aiden; Chinese, Japanese, Korean covered too).
Size breakdown (binary GiB/MiB)
(SI-decimal totals — what HF Hub displays: 4.52 GB / 2.74 GB / 1.93 GB.)
[^1]: Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice — upstream full-precision reference. [^2]: mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit — community INT8 MLX build (no 4bit upstream variant).
Recipe
# 1. LLM: bf16 -> int4 g64
python -m mlx_audio.convert \
--hf-path Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
--mlx-path ./bundle \
-q --q-bits 4 --q-group-size 64 --model-domain tts
# 2. Vocoder: fp32 -> fp16 in place
python -c "
from safetensors.numpy import load_file, save_file
import numpy as np
p = 'bundle/speech_tokenizer/model.safetensors'
sd = load_file(p)
save_file({k: v.astype(np.float16) if v.dtype == np.float32 else v for k,v in sd.items()}, p)
"Usage
from mlx_audio.tts.generate import generate_audio
generate_audio(
model="RumiLabs/Qwen3-TTS-12Hz-1.7B-CustomVoice-MLX-4bit",
text="I just heard the most amazing news.",
voice="Aiden", # English male preset
instruct="speaking with bright, lively energy at a faster pace.",
temperature=0.3, # critical — temp=0 truncates 2/5 emotions
output_path=".", file_prefix="out", audio_format="wav",
)English speakers: Aiden (Sunny American male, clear midrange) — current default. Ryan (Dynamic male, strong rhythmic drive) — alternative.
`temperature` matters. At temperature=0.0 (greedy), the LLM hits EOS prematurely on some emotion templates, producing audio truncated to "I just..." with seconds of silence padding. temperature=0.3 resolves this without quality cost.
What does NOT work (negative results worth documenting)
We attempted further INT4-g64 quantization of the non-LLM weights (the 53% of model.safetensors left at BF16 by mlx-audio's model_quant_predicate). All failed:
The mlx-audio model_quant_predicate skip list in qwen3_tts.py:241-249 is load-bearing, not conservative defaults. Whoever wrote that predicate already ran this experiment. The 1.80 GiB / 5.24 GiB-peak bundle is at the floor for this codepath.
Remaining real levers: streaming Code2Wav inference (multi-day engineering, 500 MB–1 GB peak RAM savings, no disk impact), or vocoder retrain at lower precision (multi-week).
MLX-Swift port — runs natively on iPhone (2026-06-04)
The complete stack in this repo (1.7B talker, 16-codebook code predictor, FP16 Code2Wav vocoder) is ported to MLX-Swift and validated stage-by-stage against the Python-MLX reference (mlx_audio qwen3_tts):
- Talker: teacher-forced parity at the measured bf16 near-tie floor (c0 agreement 0.98–1.00); end-to-end acceptance — Swift codec streams → reference vocoder → whisper-large-v3-turbo WER 0.000.
- Vocoder: SNR 55–58 dB vs reference wavs (the PCM16 floor of the goldens), exact sample counts, whisper WER 0.000 across 5 emotion instructs.
Sources: `swift/tts/` in the rumi-cerberus release repo (TTSTalker / TTSSession / TTSVocoder), mirrored below under swift/ for self-containment.
Measured on iPhone 16 Pro Max (rumi-cerberus ↔ TTS sequential swap)
All on-device-synthesized utterances transcribe exactly (whisper-large-v3-turbo). Deployment note: decode the vocoder in chunks (15 frames + 5 context — see TTSVocoder.decodeChunked); the single-shot graph's transient working set is jetsam-fatal on iOS even though steady-state memory fits. Sampling: voice="Aiden", temp=0.3 (temp 0 occasionally truncates on sad-style instructs), topk 50, repetitionpenalty 1.05.
Streaming voice-out update (2026-06-06)
The production loop now streams: codec frames vocode in 15-frame chunks onto an AVAudioPlayerNode while the talker keeps generating, and the LLM restores during the playback tail. Measured on the same device:
Generation is slower than real-time (audio/compute RTF 0.73–0.87, thermal-dependent), so the player is fed through a closed-loop prebuffer (EMA of measured RTF × token-estimated reply length); if a very long reply outruns the buffer, playback takes a single pause-and-rebuffer (node-clock detected) rather than accumulating stutters. Hook: generateSampled(..., onFrame:) in swift/TTSSession.swift. Release and debug builds measure identically (decode 17.1–17.3 tok/s, RTF 0.78–0.84 both) — the compute is in Metal library kernels.
License
Apache-2.0, inherited from Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice.
