CoolFace
Modelpublic

aufklarer/Magpie-TTS-Multilingual-357M-MLX-4bit

sourceHugging Faceotherupdated 4mo agoView on Hugging Face
0likes76downloads
Model Card

Magpie-TTS-Multilingual-357M-MLX-4bit

MLX port of NVIDIA Magpie-TTS Multilingual 357M, an autoregressive multi-codebook TTS model over the Nano-Codec 22 kHz / 1.89 kbps / 21.5 fps vocoder, quantized to INT4 weight-only for Apple Silicon.

Recommended for on-device deployment. ~34% the size of FP16. Round-trip ASR CER 0% on the en/de/zh samples we tested.

Model

Total parameters357 M (text encoder 99 M + decoder 90 M + LocalTransformer 1 M + NanoCodec 62 M + audio embeddings)
ArchitectureCausal Transformer encoder (6L, d=768) + Causal Transformer decoder (12L, d=768) + LocalTransformer codebook AR (1L, d=256) + Causal HiFi-GAN decoder
Audio8 codebooks × 2024 codes, 22.05 kHz mono, 21.5 fps
LanguagesEN, ES, DE, FR, IT, VI, ZH, HI, JA
Speakers5 baked (John, Sofia, Aria, Jason, Leo)
Bundle size245 MB on disk
Layout4-bundle MLX (textencoder / decoderprefill / decoderstep / nanocodecdecoder)

Files

FileSizeDescription
text_encoder/model.safetensors57.7 MBtext encoder weights (INT4)
decoder_prefill/model.safetensors68.6 MBdecoder prefill weights (INT4)
decoder_step/model.safetensors68.6 MBdecoder step weights (INT4)
nanocodec_decoder/model.safetensors63.2 MBnanocodec decoder weights (INT4)
tokenizer/*.json~30 KB eachper-language tokenizer config (8 langs + manifest)
manifest.json<1 KBSHA256 + sizes manifest

The 4-bundle layout splits the model into:

  • text_encoder — runs once per utterance over the phoneme sequence
  • decoder_prefill — batch-prefills the 110-step baked speaker context into the KV cache (~10× faster than a sequential cold start)
  • decoder_step — single AR step over the next audio frame; shares weights with decoder_prefill
  • nanocodec_decoder — codes → 22.05 kHz waveform (always FP16; per FluidInference's data, quantizing the codec yields no runtime savings)

Round-trip validation

End-to-end TTS → faster-whisper large-v3 ASR on a held-out sentence per language (Character Error Rate):

Languageenesdefritvizhhi
CER0.00%0.00%0.00%0.00%0.00%<8% (tone)<2% (1 added interjection)mixed-script Whisper artifact

Usage

python
import json
from pathlib import Path
import mlx.core as mx

# 1. Tokenize text in your app (Swift) — see speech-swift's KokoroTTS
#    pattern. For Japanese, use Apple's CFStringTokenizer + katakana → IPA.
# 2. Load the 3 sub-models and run the AR loop.
from huggingface_hub import snapshot_download
bundle = Path(snapshot_download("aufklarer/Magpie-TTS-Multilingual-357M-MLX-4bit"))

# Production usage: see https://github.com/soniqo/speech-swift.

The production Swift integration handles tokenization, the AR loop, KV-cache management, and audio rendering. This HuggingFace bundle exists for researchers and SDK developers building atop the MLX weights directly.

Source

License

NVIDIA Open Model License — inherited from upstream Magpie-TTS Multilingual. Suitable for commercial use; please review the license text linked above.