CoolFace
Modelpublic

tpsjr7/acestep-captioner-GGUF

sourceHugging Facemitupdated 14d agoView on Hugging Face
2likes253downloads
Model Card

ACE-Step Captioner GGUF

This repository contains a llama.cpp-compatible GGUF conversion of ACE-Step/acestep-captioner.

The initial upload is intended to include:

  • —acestep-captioner-Q4_K_M.gguf
  • —acestep-captioner-Q6_K.gguf
  • —acestep-captioner-Q8_0.gguf
  • —acestep-captioner-mmproj-bf16.gguf

Files

  • —acestep-captioner-Q4_K_M.gguf: Quantized text model for inference with llama.cpp
  • —acestep-captioner-Q6_K.gguf: Higher-quality 6-bit quantized text model
  • —acestep-captioner-Q8_0.gguf: Higher-quality 8-bit quantized text model
  • —acestep-captioner-mmproj-bf16.gguf: Multimodal projector required for audio input

llama.cpp

This model requires a recent llama.cpp build with Qwen2.5-Omni audio support.

Tested with this fork / branch which fixed audio inference bugs: https://github.com/tpsjr7/llama.cpp/tree/ted/fix-qwen-audio-cleanup-merge

Example:

bash
llama-cli -m ./acestep-captioner-Q4_K_M.gguf \
  --mmproj ./acestep-captioner-mmproj-bf16.gguf \
  --audio ./song.mp3 \
  -p "*Task* Describe this audio in detail" \
  -n 512 --temp 0 --single-turn --simple-io -ngl 999 --ctx-size 8192

Swap in acestep-captioner-Q6_K.gguf or acestep-captioner-Q8_0.gguf if you want a less aggressive quantization.