CoolFace
Modelpublic

saxyZ/audio.cpp-gguf-fork

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes3.2kdownloads
Model Card

audio.cpp GGUF Model Packages

This directory contains audio.cpp-native GGUF conversions of multiple speech models. These files are intended for use with audio.cpp.

For conversion details, supported layouts, direct-file loading, sidecar embedding, and the latest compatibility notes, see the audio.cpp GGUF guide:

  • https://github.com/0xShug0/audio.cpp/blob/main/docs/gguf.md

!!! Automated audio checks are intentionally strict and may flag length, log-mel, or transcript drift that can still sound acceptable to human listeners. Validate the exact file, backend, and route you plan to use. These converted weights are provided as-is; use them at your own risk.

Files

Tested summarizes the current audio.cpp path-test status. See the GGUF guide above for the full matrix and drift notes.

DirectoryFilesaudio.cpp familyTestedOriginal model license
ACE-Step1.5-GGUFbase/ace-step-1.5-base-bf16.gguf, base/ace-step-1.5-base-q8_0.gguf, turbo/ace-step-1.5-turbo-bf16.gguf, turbo/ace-step-1.5-turbo-q8_0.gguface_step16-bit + Q8 driftMIT
BS-RoFormer-ep368-GGUFbs-roformer-ep368-q8_0.ggufbs_roformerQ8 passApache-2.0
Chatterbox-GGUFchatterbox-f16.gguf, chatterbox-q8_0.ggufchatterbox16-bit + Q8 ASR-match driftMIT
Citrinet-ASR-GGUFcitrinet-asr-q8_0.ggufcitrinet_asrQ8 passCC-BY-4.0
Confucius4-TTS-GGUFconfucius4-tts-orig.ggufconfucius4_ttsorig passApache-2.0
DotTTS-MF-GGUFdots-tts-mf-bf16.ggufdots_ttsexperimentalApache-2.0
DotTTS-SOAR-GGUFdots-tts-soar-orig.gguf, dots-tts-soar-bf16.ggufdots_ttsexperimentalApache-2.0
DramaBox-GGUFdramabox-q8_0.ggufdramaboxQ8 passLTX-2 Community License
Fish-Audio-S2-Pro-GGUFfish-audio-s2-pro-bf16.gguf, fish-audio-s2-pro-q8_0.gguffish_audio16-bit + Q8 passFish Audio Research License
Fun-ASR-Nano-2512-GGUFfun-asr-nano-2512-f16.gguf, fun-asr-nano-2512-q8_0.gguffun_asr_nano16-bit + Q8 passFunASR Model Open Source License Agreement v1.1
HeartMuLa-GGUFheartmula-f16.gguf, heartmula-q8_0.ggufheartmula16-bit + Q8 driftApache-2.0
HTDemucs-GGUFhtdemucs-f16.gguf, htdemucs-q8_0.ggufhtdemucs16-bit pass, Q8 driftMIT
Higgs-Audio-v3-STT-GGUFhiggs-audio-v3-stt-f16.gguf, higgs-audio-v3-stt-q8_0.ggufhiggs_audio_stt16-bit + Q8 passApache-2.0
Higgs-Audio-v3-TTS-4B-GGUFhiggs-audio-v3-tts-4b-bf16.gguf, higgs-audio-v3-tts-4b-q8_0.ggufhiggs_audio_tts16-bit + Q8 passBoson Higgs TTS 3 Research and Non-Commercial License
Hviske-v5.3-GGUFhviske-v5.3-q8_0.ggufhviske_asrQ8 passCC-BY-NC-4.0
IndexTTS2-GGUFindex-tts2-orig.gguf, index-tts2-f16.gguf, index-tts2-q8_0.ggufindex_tts2orig + 16-bit pass/drift, Q8 ASR-match driftbilibili Model Use License Agreement
IndexTTS2.5-GGUFindex-tts2_5-orig.gguf, index-tts2_5-f16.gguf, index-tts2_5-q8_0.ggufindex_tts2self-contained smoke pass; Q8/F16/orig CUDA load + synthesizebilibili Model Use License Agreement
Inflect-Micro-v2-GGUFinflect-micro-v2-orig.ggufinflect_v2orig passApache-2.0
Irodori-TTS-500M-v3-GGUFirodori-tts-500m-v3-f16.gguf, irodori-tts-500m-v3-q8_0.ggufirodori_tts16-bit pass, Q8 driftMIT
Irodori-TTS-600M-v3-VoiceDesign-GGUFirodori-tts-600m-v3-voicedesign-f16.gguf, irodori-tts-600m-v3-voicedesign-q8_0.ggufirodori_tts16-bit pass, Q8 driftMIT
Irodori-TTS-v4-Small-GGUFirodori-tts-v4-small-f16.gguf, irodori-tts-v4-small-q8_0.ggufirodori_ttsv4.1 checkpoint, 16-bit + Q8 passMIT
Kroko-ASR-GGUFkroko-en-community-64-l-q8_0.ggufkroko_asrQ8 passCC-BY-SA community model license
MOSS-TTS-Local-v1.5-GGUFmoss-tts-local-v1.5-bf16.gguf, moss-tts-local-v1.5-q8_0.ggufmoss_tts_local16-bit pass, Q8 ASR-match driftApache-2.0
MOSS-TTS-Nano-100M-GGUFmoss-tts-nano-100m-bf16.gguf, moss-tts-nano-100m-q8_0.ggufmoss_tts_nano16-bit pass, Q8 ASR-match driftApache-2.0
MagpieTTS-Multilingual-357M-GGUFmagpie-tts-multilingual-357m-orig.ggufmagpie_ttsexperimentalNVIDIA Open Model License
Mel-Band-RoFormer-GGUFmel-band-roformer-f16.gguf, mel-band-roformer-q8_0.ggufmel_band_roformer16-bit + Q8 driftMIT
MioCodec-25Hz-44.1kHz-v2-GGUFmiocodec-25hz-44khz-v2-orig.gguf, miocodec-25hz-44khz-v2-f16.gguf, miocodec-25hz-44khz-v2-q8_0.ggufmiocodecorig pass, 16-bit + Q8 driftMIT
MioTTS-1.7B-GGUFmiotts-1.7b-orig.gguf, miotts-1.7b-bf16.gguf, miotts-1.7b-q8_0.ggufmiottsorig pass, 16-bit drift, Q8 ASR-match driftApache-2.0
MuScriptor-Small-GGUFmuscriptor-small-f32.ggufmuscriptorF32 passCC-BY-NC-4.0
Nemotron-3.5-ASR-Streaming-0.6B-GGUFnemotron-3.5-asr-streaming-0.6b-f16.gguf, nemotron-3.5-asr-streaming-0.6b-q8_0.ggufnemotron_asr16-bit pass, Q8 minor filler driftOpenMDW-1.1
OmniVoice-GGUFomnivoice-bf16.gguf, omnivoice-f16.gguf, omnivoice-q8_0.ggufomnivoice16-bit + Q8 driftApache-2.0
Parakeet-TDT-0.6B-v3-GGUFparakeet-tdt-0.6b-v3-f16.gguf, parakeet-tdt-0.6b-v3-q8_0.ggufparakeet_tdt16-bit + Q8 passCC-BY-4.0
PocketTTS-GGUFenglish/, german/, italian/, portuguese/, spanish/ each contain bf16 and q8_0 GGUFspocket_tts16-bit pass, Q8 driftCC-BY-4.0
Qwen3-ASR-0.6B-GGUFqwen3-asr-0.6b-f16.gguf, qwen3-asr-0.6b-q8_0.ggufqwen3_asr16-bit + Q8 passApache-2.0
Qwen3-ASR-1.7B-GGUFqwen3-asr-1.7b-f16.gguf, qwen3-asr-1.7b-q8_0.ggufqwen3_asr16-bit + Q8 passApache-2.0
Qwen3-ForcedAligner-0.6B-GGUFqwen3-forced-aligner-0.6b-f16.gguf, qwen3-forced-aligner-0.6b-q8_0.ggufqwen3_forced_aligner16-bit + Q8 passApache-2.0
Qwen3-TTS-12Hz-0.6B-Base-GGUFqwen3-tts-12hz-0.6b-base-bf16.gguf, qwen3-tts-12hz-0.6b-base-q8_0.ggufqwen3_ttsBF16 exact vs safetensors, Q8 ASR-match driftApache-2.0
Qwen3-TTS-12Hz-1.7B-Base-GGUFqwen3-tts-12hz-1.7b-base-orig.gguf, qwen3-tts-12hz-1.7b-base-bf16.gguf, qwen3-tts-12hz-1.7b-base-q8_0_v2.ggufqwen3_ttsorig pass, 16-bit + Q8 ASR-match driftApache-2.0
Qwen3-TTS-12Hz-1.7B-CustomVoice-GGUFqwen3-tts-12hz-1.7b-customvoice-bf16.gguf, qwen3-tts-12hz-1.7b-customvoice-q8_0.ggufqwen3_tts16-bit + Q8 ASR-match driftApache-2.0
Qwen3-TTS-12Hz-1.7B-VoiceDesign-GGUFqwen3-tts-12hz-1.7b-voicedesign-bf16.gguf, qwen3-tts-12hz-1.7b-voicedesign-q8_0.ggufqwen3_tts16-bit + Q8 ASR-match driftApache-2.0
RVC-GGUFrvc-f16.ggufrvcF16 passMIT
SeedVC-MLX-GGUFseed-vc-mlx-orig.gguf, seed-vc-mlx-f16.gguf, seed-vc-mlx-q8_0.ggufseed_vc16-bit + Q8 driftGPL-3.0
Sortformer-Diar-4spk-v1-GGUFsortformer-diar-4spk-v1-f16.gguf, sortformer-diar-4spk-v1-q8_0.ggufsortformer_diar16-bit + Q8 passCC-BY-NC-4.0
Stable-Audio-3-Medium-GGUFstable-audio-3-medium-f16.gguf, stable-audio-3-medium-q8_0.ggufstable_audio16-bit + Q8 driftStability AI Community License
Stable-Audio-3-Small-Music-GGUFstable-audio-3-small-music-f16.gguf, stable-audio-3-small-music-q8_0.ggufstable_audio16-bit + Q8 driftStability AI Community License
Stable-Audio-3-Small-SFX-GGUFstable-audio-3-small-sfx-f16.gguf, stable-audio-3-small-sfx-q8_0.ggufstable_audio16-bit + Q8 driftStability AI Community License
Supertonic-3-GGUFsupertonic-3-orig.gguf, supertonic-3-f16.gguf, supertonic-3-q8_0.ggufsupertonicF32/orig pass; f16 not tested; Q8 unsupported dtypeBigScience Open RAIL-M
Vevo2-GGUFvevo2-orig.gguf, vevo2-f16.gguf, vevo2-q8_0.ggufvevo2orig + 16-bit pass/drift; Q8 mixed route driftCC-BY-NC-ND-4.0
VibeVoice-1.5B-GGUFvibevoice-1.5b-bf16.gguf, vibevoice-1.5b-q8_0.gguf, vibevoice-1.5b-q4-ios.ggufvibevoice16-bit pass, Q8 driftMIT
VibeVoice-ASR-GGUFvibevoice-asr-f16.gguf, vibevoice-asr-q8_0.ggufvibevoice_asr16-bit + Q8 passMIT
VoxCPM2-GGUFvoxcpm2-orig.gguf, voxcpm2-bf16.gguf, voxcpm2-q8_0.ggufvoxcpm2orig pass, 16-bit + Q8 ASR-match driftApache-2.0
Voxtral-Mini-4B-Realtime-2602-GGUFvoxtral-mini-4b-realtime-2602-bf16.gguf, voxtral-mini-4b-realtime-2602-q8_0.gguf, voxtral-mini-4b-realtime-2602-q4_k.ggufvoxtral_realtime16-bit + Q8 pass; Q4_K quick check passedApache-2.0

Q8 Notes

  • Chatterbox Q8 is intentionally mixed type. Graph-sensitive scalar, norm, bias, and side tensors stay in non-Q8 types while matmul-compatible weights are quantized.
  • PocketTTS Q8 keeps the four flow_lm.flow_net.time_embed.*.mlp.{0,2}.weight tensors in Q8 in addition to the default converter selection. conditioner.embed, cond_embed, and Mimi conv tensors are not forced to Q8 because tested outputs drifted or the current conv path casts quantized conv weights back to F32.
  • Voxtral Q4K is smaller than Q80 and was faster in a quick CUDA path check, with transcripts matching Q8_0 except for one capitalization-only difference.

Usage

Pass a GGUF file directly as --model:

bash
audiocpp_cli --task tts --family supertonic --model Supertonic-3-GGUF/supertonic-3-orig.gguf --backend cuda --language en --text "Hello." --voice-id M1 --out out.wav

For ASR:

bash
audiocpp_cli --task asr --family qwen3_asr --model Qwen3-ASR-0.6B-GGUF/qwen3-asr-0.6b-f16.gguf --backend cuda --audio speech.wav --text "" --text-out transcript.txt

License

Each GGUF file is a converted form of its original model. Use and redistribution are governed by the corresponding original model license listed above. Please review the original model card and license terms before using or redistributing any converted weights.