tpsjr7/acestep-captioner-GGUF
2253
ACE-Step Captioner GGUF
This repository contains a llama.cpp-compatible GGUF conversion of ACE-Step/acestep-captioner.
The initial upload is intended to include:
acestep-captioner-Q4_K_M.ggufacestep-captioner-Q6_K.ggufacestep-captioner-Q8_0.ggufacestep-captioner-mmproj-bf16.gguf
Files
acestep-captioner-Q4_K_M.gguf: Quantized text model for inference withllama.cppacestep-captioner-Q6_K.gguf: Higher-quality 6-bit quantized text modelacestep-captioner-Q8_0.gguf: Higher-quality 8-bit quantized text modelacestep-captioner-mmproj-bf16.gguf: Multimodal projector required for audio input
llama.cpp
This model requires a recent llama.cpp build with Qwen2.5-Omni audio support.
Tested with this fork / branch which fixed audio inference bugs: https://github.com/tpsjr7/llama.cpp/tree/ted/fix-qwen-audio-cleanup-merge
Example:
llama-cli -m ./acestep-captioner-Q4_K_M.gguf \
--mmproj ./acestep-captioner-mmproj-bf16.gguf \
--audio ./song.mp3 \
-p "*Task* Describe this audio in detail" \
-n 512 --temp 0 --single-turn --simple-io -ngl 999 --ctx-size 8192Swap in acestep-captioner-Q6_K.gguf or acestep-captioner-Q8_0.gguf if you want a less aggressive quantization.
