audio-cpp/MiniMax-Music3-GGUF
MiniMax Music 3 GGUF
GGUF package for MiniMax Music 3 for audio.cpp.
Star our repo so you don't miss important updates! https://github.com/0xShug0/audio.cpp
Upstream model: https://huggingface.co/MiniMaxAI/MiniMax-Music3 Upstream license: https://huggingface.co/MiniMaxAI/MiniMax-Music3/blob/main/LICENSE
Compatible Q8 GGUF package: https://huggingface.co/joemattie/MiniMax-Music3-GGUF
Notes
- The implementation is available on the
mainbranch and release 0.6.1. - The current runtime uses model-local resource loading instead of treating the v1 spec as the runtime contract. This keeps component selection flexible while the package layout and option surface settle.
- The default component mix favors Q40 for the large language model and flow transformer, with Q80 for the RVQ depth decoder.
- BF16, Q80, and Q40 component variants are included for quality/performance comparison.
- Longer generations such as five-minute songs are supported as long-form runs, but they are currently tuned for completion and quality checks rather than realtime throughput.
- Memory usage remains an active optimization target for larger durations and alternate component mixes.
Performance Snapshot
Measured on an RTX 5090 with CUDA using a 30-second lyric generation request, 30 flow steps, flow guidance scale 1.7, AR guidance scale 1.5, and top-k 50. Peak VRAM is the observed nvidia-smi process peak during a warmup-plus-measured-request run.
Quick Start
The prompt is for smoke end-to-end test only.
audiocpp_cli \
--task gen \
--family minimax_music3 \
--model MiniMax-Music3-GGUF \
--backend cuda \
--text "A bright pop rock song with clean drums and a clear male vocal." \
--request-option 'lyrics=City lights are shining low. I keep moving with the glow. Turn it up and let it fly. Sing the melody tonight.' \
--request-option duration_sec=10 \
--request-option num_inference_steps=30 \
--out output.wavComponents
Default audio.cpp component mix:
language_model_q4_0.ggufrvq_depth_decoder_q8_0.gguftransformer_q4_0.ggufcondition_encoder.ggufvocoder.gguf
The BF16, Q80, and Q40 component variants are included for measurement and quality/performance comparison.
Component GGUFs can be selected explicitly for experiments:
--session-option minimax_music3.language_model_gguf=language_model_bf16.gguf
--session-option minimax_music3.language_model_gguf=language_model_q8_0.gguf
--session-option minimax_music3.rvq_depth_decoder_gguf=rvq_depth_decoder_bf16.gguf
--session-option minimax_music3.rvq_depth_decoder_gguf=rvq_depth_decoder_q8_0.gguf
--session-option minimax_music3.flow_transformer_gguf=transformer_bf16.gguf
--session-option minimax_music3.flow_transformer_gguf=transformer_q8_0.ggufLicense
This GGUF package follows the upstream MiniMax-Music3 COMMUNITY LICENSE. The license text is included in LICENSE; review it before use, especially for commercial deployment.
