axiom-of-choice/qwen3-1.7b-es-reasoning-gguf
qwen3-1.7b-es-reasoning-gguf
GGUF quantizations for llama.cpp, Ollama and LM Studio. Quantized from a local merge of `axiom-of-choice/qwen3-1.7b-es-reasoning-peft` onto Qwen/Qwen3-1.7B.
Files
Q4_K_MQ8_0
Q4_K_M is the default most tools pull -- smallest with negligible quality loss for most use. Q8_0 is near-lossless if you have the RAM/VRAM for it.
Usage
ollama run hf.co/axiom-of-choice/qwen3-1.7b-es-reasoning-gguf:Q4_K_M
# or
llama-cli -hf axiom-of-choice/qwen3-1.7b-es-reasoning-gguf:Q4_K_M -p "¿Cuánto es 17 por 24?"What the numbers below do NOT cover
The merge-verification table is measured on the bf16 merge, before this GGUF's own quantization step. Q4/Q8 quantization is a further lossy step on top of that and is not separately measured here -- treat it as an additional, unquantified source of drift on top of the number below, not as covered by it.
Did merging (pre-quantization) change the model?
This is checked against the unmerged PEFT adapter's own generations (results/peft_parity/qwen3-1.7b-s6.0-n100.json), not the original MLX numbers -- the question here is only "did merging change the model", and MLX-vs-PyTorch divergence is a separate, already-documented story on the adapter repo's card.
Paired counts: 2 correct only here, 0 correct only unmerged, out of 20. This drifted more than expected for a same-engine comparison (unlike the MLX-vs-PyTorch phase-2 numbers, there is no expected source of divergence between an adapter and its own merge). Flagged here rather than smoothed over; treat the merge as unverified until this is understood.
Training
Same recipe, dataset and evaluation as the adapter -- see `axiom-of-choice/qwen3-1.7b-es-reasoning-peft`.
