Q8_0 GGUF quantization of the FastContext 1.0 4B SFT model, ready to run with llama.cpp, Ollama, and other GGUF runtimes.
FastContext-1.0-4B-SFT-Q8_0.gguf
# llama.cpp llama-cli -hf dstolf/FastContext-1.0-4B-SFT-Q8_0-GGUF # Ollama ollama run hf.co/dstolf/FastContext-1.0-4B-SFT-Q8_0-GGUF