musk12/apple-fastvlm-qwen2-q4-gguf
016
FastVLM Qwen2 (Q4KM, GGUF)
A Qwen2 GGUF model quantized in Q4KM format for CPU memory-efficient inference. This FastVLM LLM backbone is optimized for low-memory and low-latency CPU deployment.
Model file
fastvlm_qwen2_q4km.gguf
Usage
Use this model with llama.cpp-compatible runtimes for fast CPU-based language generation in FastVLM.
./llama-cli -m fastvlm_qwen2_q4km.gguf -p "Your prompt here"Base models
Derived from:
