CoolFace
Modelpublic

ibarron/MOSS-TTS-v1.5-Q6_K-GGUF

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
1likes115downloads
Model Card

MOSS-TTS v1.5 — Q6_K GGUF

A Q6_K GGUF quantization of the MOSS-TTS v1.5 backbone (Qwen3-8B), packaged for the moss_tts_delay llama.cpp inference path instead of the full bf16 PyTorch weights — so MOSS-TTS fits comfortably on memory-constrained machines (e.g. 16–24 GB Apple Silicon).

Contents

  • —`backbone_q6_k.gguf` — the MOSS-TTS v1.5 backbone quantized to Q6_K, with the token-embedding tensor (`token_embd`) kept at F16. This is deliberate and important: MOSS-TTS reads per-token text/audio embeddings directly from that table, and quantizing it below F16 degrades output. Quantize everything else, keep token_embd at F16 (llama-quantize --token-embedding-type f16).
  • —`extracted/` — the npy embedding + LM-head tables and tokenizer that the moss_tts_delay llama.cpp pipeline loads alongside the GGUF (the audio/codec heads that don't live inside the backbone GGUF).

Not included (get these from the official sources)

How it was built

Convert the official v1.5 weights to F16 GGUF (convert_hf_to_gguf.py), then llama-quantize --token-embedding-type f16 … Q6_K. On an internal speech-quality battery (word-error-rate, speaker-similarity, and mel-cepstral-distortion), Q6_K matched Q8 while using less memory — making it a good default quant for running MOSS-TTS locally.

License & credit

Apache-2.0, inherited from MOSS-TTS v1.5 by the OpenMOSS team. This repository only re-packages their model as a quantized GGUF; all model credit belongs to OpenMOSS. Please cite and follow the upstream model's terms.