CoolFace
Modelpublic

PowerBeef02/Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit

sourceHugging Faceapache-2.0updated 10d agoView on Hugging Face
0likes213downloads
Model Card

PowerBeef02/Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit

Vocello production artifact. Derived from `mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit` at revision 41d3337e8b7f2843a75841595fc14e4b9a7a4b96 (itself an MLX conversion of the corresponding `Qwen/Qwen3-TTS` checkpoint), with one change: the 622 MB BF16 talker.model.text_embedding tensor is quantized to affine 8-bit (group size 64) by Vocello's pinned conversion tooling (scripts/convert_text_embedding_8bit.py, python-mlx 0.32.0). Every other tensor is byte-identical to the source revision.

Measured on the Vocello support floor (Mac mini M2 8 GB): about 278 MB less resident memory and 292 MB less download per artifact at parity real-time factor, with clean deterministic audio QC. Measurement details live in the Vocello repository (benchmarks/OPTIMIZATION.md §N).

These files are consumed by Vocello's fail-closed model catalog, which pins this repository at an exact revision with per-file SHA-256 digests.