CoolFace
Modelpublic

ngquocvinh/Spark-X2.5-1.7B-GGUF

sourceHugging Faceapache-2.0updated 14d agoView on Hugging Face
1likes2.8kdownloads
Model Card

Spark-X2.5-1.7B GGUF

Community GGUF quantizations of XHToken/Spark-X2.5-1.7B. This repository contains nine quantized files for local inference. No training or fine-tuning was performed.

<div align="center" style="background-color:#f59e0b;color:#ffffff;padding:16px 20px;border-radius:10px;line-height:1.7;"> ☕ If this GGUF made your day easier, a coffee would make mine.<br> <a href="https://ko-fi.com/ngquocvinh" style="color:#ffffff;"><strong style="color:#ffffff;">Send a coffee ☕</strong></a><br> I build and test these releases myself. Your coffee helps keep me going.<br> Thank you for supporting this work. </div>

Files

QuantizationFile size (GiB)A10M generation token/sValidationRecommendation / Notes
Q8_01.70175.54Load/generate passHigh quality.
Q6_K1.31199.07Load/generate passHigh quality.
Q5KM1.17223.04Load/generate passDaily use.
Q4KM1.03241.87Load/generate passRecommended default.
Q3KM0.87203.60Load/generate passLower-memory profile.
Q2_K0.74231.70Load/generate passAggressive low-memory profile.
IQ2_XS0.61245.03Load/generate passExperimental.
IQ1_M0.54252.02Load/generate passExperimental.
Q1_00.40336.65Load/generate passExperimental / legacy minimum-memory option.

The A10M generation figures were measured with single-stream llama-bench on an NVIDIA A10M.

Q1/Q2 and the IQ variants can lose instruction following, reasoning, and tool-call reliability. Validate the chosen file on the workload that matters to you.

License and attribution

The upstream model is released under Apache License 2.0. Preserve the upstream attribution and license when redistributing these derivative files. This is a community GGUF quantization, not an official XHToken/SparkLLM release or endorsement.

Checksums are available in `SHA256SUMS.txt`.