CoolFace
Modelpublic

infinia-ai/magpie-tts-13lang-357m

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes9downloads
Model Card

Magpie-TTS 13-Language (357M) โ€” with Swahili

๐Ÿ“„ Paper: Cross-Lingual Grafting: Extending Frozen Codec Speech Language Models to Unseen Low-Resource Languages โ€” PDF.

A single 13-language text-to-speech checkpoint: the twelve languages of NVIDIA's `magpie_tts_multilingual_357m` (Arabic, German, English, Spanish, French, Hindi, Italian, Japanese, Korean, Portuguese, Vietnamese, Chinese) plus Swahili (Kiswahili) โ€” added by a community team without regressing any of the original twelve.

Swahili was grafted onto the frozen base model with a dedicated byte-level tokenizer, a warm-started input-embedding surgery, and multilingual rehearsal. On held-out Swahili the model reaches 8.8% character error rate (median 0.0%, MMS-Swahili ASR), and the twelve base languages show no measurable regression versus the untouched base. Full method: see the accompanying paper "Cross-Lingual Grafting: Extending Frozen Codec Speech Language Models to Unseen Low-Resource Languages" (New Emerging Technologies, a subsidiary of Infinia Technologies).

โš ๏ธ Community model. Swahili support was added by the community and is not provided or endorsed by NVIDIA. The twelve base languages follow the base model.

Highlights

BaseMagpie-TTS-Multilingual 357M (Koel-TTS family)
Audio codecLow Frame-rate Speech Codec (22.05 kHz, 1.89 kbps, 21.5 fps)
Languages13 (12 base + Swahili)
Swahili quality8.8% mean / 0.0% median CER (MMS-sw)
Base regression~0 (mean recognizer CER identical to base)
Voices5 baked speakers (2 female, 3 male), selected by index
Trained on1ร— 128 GB unified-memory GPU

Usage (NVIDIA NeMo)

python
from nemo.collections.tts.models import MagpieTTSModel
import soundfile as sf, numpy as np, torch, random

m = MagpieTTSModel.from_pretrained("infinia-ai/magpie-tts-13lang-357m").eval().cuda()

# Deterministic inference: pin the seed so identical inputs give identical audio
def seed(s=1234):
    random.seed(s); np.random.seed(s); torch.manual_seed(s); torch.cuda.manual_seed_all(s)

seed(1234)
audio, alen = m.do_tts(
    "Habari ya asubuhi. Karibu kwenye jaribio la sauti ya Kiswahili.",
    language="sw",            # native Swahili code; also en, de, es, fr, it, vi,
                              # zh, hi, ja, pt-BR, ko, ar-MSA, ar-SA, ar-AE
    speaker_index=0,          # 0 Aria(F) 1 Jason(M) 2 John(M) 3 Leo(M) 4 Sofia(F)
    apply_TN=False, use_cfg=True,
)
w = np.squeeze(audio.detach().float().cpu().numpy())
sf.write("out.wav", (w[0] if w.ndim > 1 else w).astype(np.float32), 22050)

Determinism. The released tokenizers use phoneme_probability=1.0 (deterministic phonemization) and the sampler defaults (temperature 0.7, top-k 80, CFG 2.5) match the base model. With a fixed RNG seed, identical (text, language, voice, CFG, seed) inputs produce byte-identical audio; change the seed for a different rendering. For the occasional degenerate silent generation, resample at a new seed (retry-on-silence).

Language codes

en, de, es, fr, it, vi, zh, hi, ja, pt-BR, ko, ar-MSA, ar-SA, ar-AE, sw.

Training data

  • โ€”Swahili: Bateesa/kiswahili-tts-dataset (CC-BY) + FLEURS sw_ke (CC-BY), ~27.8 h, resampled to 22.05 kHz.
  • โ€”Rehearsal (12 base languages): FLEURS, ~500 utterances/language, each routed to its native tokenizer.

Please cite FLEURS (Conneau et al., arXiv:2205.12446) and the Koel-TTS (arXiv:2502.05236) and Low Frame-rate Speech Codec (arXiv:2409.12117) papers.

Intended use & limitations

Research and product speech synthesis for the 13 supported languages. Out of scope: impersonation of real individuals, deceptive or harmful synthetic media. The model synthesizes a fixed set of 5 baked voices; it does not support zero-shot voice cloning (the base model's context encoder is not included). Swahili data skews toward read/literary speech and one dominant speaker; FLEURS adds multi-speaker breadth at upsampled 16 kHz. 357M parameters bound absolute quality.

Responsible use

Generated audio carries the base model's watermark and synthetic-speech disclosure. Do not use to deceive. Disclose that audio is AI-generated.

License & attribution

Released under the NVIDIA Open Model License (see LICENSE and NOTICE). Derived from nvidia/magpie_tts_multilingual_357m. "Licensed by NVIDIA Corporation under the NVIDIA Open Model License." Training data: FLEURS (CC-BY), Bateesa Kiswahili TTS (CC-BY).