CoolFace
Modelpublic

abenzerps/Spark-X2.5-4B-GGUF

sourceHugging Faceapache-2.0updated 16d agoView on Hugging Face
22likes153kdownloads
Model Card
[!IMPORTANT] Compatibility: These GGUF files require llama.cpp b10828 or later, which includes official support for the Spark-X2.5 (spark2_5) architecture. Applications with a bundled runtime must use an equivalent or newer build. llama.cpp support

Spark-X2.5-4B GGUF

GGUF quantizations of XHToken/Spark-X2.5-4B, a 4B general-purpose language model for reasoning, coding, tool use, and agentic workflows. Native context: 1,048,576 tokens (1M).

Benchmarks

[image]

Benchmark results reported by XHToken for Spark-X2.5-4B in thinking mode.

GGUF files

QuantizationFileSize
Q4_0Spark-X2.5-4B-Q4_0.gguf2.41 GB
Q4KMSpark-X2.5-4B-Q4_K_M.gguf2.60 GB
Q5KMSpark-X2.5-4B-Q5_K_M.gguf2.98 GB
Q6_KSpark-X2.5-4B-Q6_K.gguf3.38 GB
Q8_0Spark-X2.5-4B-Q8_0.gguf4.38 GB

Includes the upstream chat_template.jinja. Checksums: SHA256SUMS.txt.

Usage

bash
llama-cli -m Spark-X2.5-4B-Q4_K_M.gguf -c 131072 -cnv

Source

  • Model: XHToken/Spark-X2.5-4B
  • Revision: ea14618d20e76b5b093d3ee20a5b9d733bb12410
  • License: Apache-2.0