infinia-ai/magpie-tts-13lang-357m
Magpie-TTS 13-Language (357M) โ with Swahili
๐ Paper: Cross-Lingual Grafting: Extending Frozen Codec Speech Language Models to Unseen Low-Resource Languages โ PDF.
A single 13-language text-to-speech checkpoint: the twelve languages of NVIDIA's `magpie_tts_multilingual_357m` (Arabic, German, English, Spanish, French, Hindi, Italian, Japanese, Korean, Portuguese, Vietnamese, Chinese) plus Swahili (Kiswahili) โ added by a community team without regressing any of the original twelve.
Swahili was grafted onto the frozen base model with a dedicated byte-level tokenizer, a warm-started input-embedding surgery, and multilingual rehearsal. On held-out Swahili the model reaches 8.8% character error rate (median 0.0%, MMS-Swahili ASR), and the twelve base languages show no measurable regression versus the untouched base. Full method: see the accompanying paper "Cross-Lingual Grafting: Extending Frozen Codec Speech Language Models to Unseen Low-Resource Languages" (New Emerging Technologies, a subsidiary of Infinia Technologies).
โ ๏ธ Community model. Swahili support was added by the community and is not provided or endorsed by NVIDIA. The twelve base languages follow the base model.
Highlights
Usage (NVIDIA NeMo)
from nemo.collections.tts.models import MagpieTTSModel
import soundfile as sf, numpy as np, torch, random
m = MagpieTTSModel.from_pretrained("infinia-ai/magpie-tts-13lang-357m").eval().cuda()
# Deterministic inference: pin the seed so identical inputs give identical audio
def seed(s=1234):
random.seed(s); np.random.seed(s); torch.manual_seed(s); torch.cuda.manual_seed_all(s)
seed(1234)
audio, alen = m.do_tts(
"Habari ya asubuhi. Karibu kwenye jaribio la sauti ya Kiswahili.",
language="sw", # native Swahili code; also en, de, es, fr, it, vi,
# zh, hi, ja, pt-BR, ko, ar-MSA, ar-SA, ar-AE
speaker_index=0, # 0 Aria(F) 1 Jason(M) 2 John(M) 3 Leo(M) 4 Sofia(F)
apply_TN=False, use_cfg=True,
)
w = np.squeeze(audio.detach().float().cpu().numpy())
sf.write("out.wav", (w[0] if w.ndim > 1 else w).astype(np.float32), 22050)Determinism. The released tokenizers use phoneme_probability=1.0 (deterministic phonemization) and the sampler defaults (temperature 0.7, top-k 80, CFG 2.5) match the base model. With a fixed RNG seed, identical (text, language, voice, CFG, seed) inputs produce byte-identical audio; change the seed for a different rendering. For the occasional degenerate silent generation, resample at a new seed (retry-on-silence).
Language codes
en, de, es, fr, it, vi, zh, hi, ja, pt-BR, ko, ar-MSA, ar-SA, ar-AE, sw.
Training data
- Swahili: Bateesa/kiswahili-tts-dataset (CC-BY) + FLEURS
sw_ke(CC-BY), ~27.8 h, resampled to 22.05 kHz. - Rehearsal (12 base languages): FLEURS, ~500 utterances/language, each routed to its native tokenizer.
Please cite FLEURS (Conneau et al., arXiv:2205.12446) and the Koel-TTS (arXiv:2502.05236) and Low Frame-rate Speech Codec (arXiv:2409.12117) papers.
Intended use & limitations
Research and product speech synthesis for the 13 supported languages. Out of scope: impersonation of real individuals, deceptive or harmful synthetic media. The model synthesizes a fixed set of 5 baked voices; it does not support zero-shot voice cloning (the base model's context encoder is not included). Swahili data skews toward read/literary speech and one dominant speaker; FLEURS adds multi-speaker breadth at upsampled 16 kHz. 357M parameters bound absolute quality.
Responsible use
Generated audio carries the base model's watermark and synthetic-speech disclosure. Do not use to deceive. Disclose that audio is AI-generated.
License & attribution
Released under the NVIDIA Open Model License (see LICENSE and NOTICE). Derived from nvidia/magpie_tts_multilingual_357m. "Licensed by NVIDIA Corporation under the NVIDIA Open Model License." Training data: FLEURS (CC-BY), Bateesa Kiswahili TTS (CC-BY).
