PowerBeef02/Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit
PowerBeef02/Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit
Vocello production artifact. Derived from `mlx-community/Qwen3-TTS-12Hz-1.7B-CustomVoice-8bit` at revision 41d3337e8b7f2843a75841595fc14e4b9a7a4b96 (itself an MLX conversion of the corresponding `Qwen/Qwen3-TTS` checkpoint), with one change: the 622 MB BF16 talker.model.text_embedding tensor is quantized to affine 8-bit (group size 64) by Vocello's pinned conversion tooling (scripts/convert_text_embedding_8bit.py, python-mlx 0.32.0). Every other tensor is byte-identical to the source revision.
Measured on the Vocello support floor (Mac mini M2 8 GB): about 278 MB less resident memory and 292 MB less download per artifact at parity real-time factor, with clean deterministic audio QC. Measurement details live in the Vocello repository (benchmarks/OPTIMIZATION.md §N).
These files are consumed by Vocello's fail-closed model catalog, which pins this repository at an exact revision with per-file SHA-256 digests.
