CoolFace
Modelpublic

worthdoing/Qwen2.5-Coder-7B-Instruct-GGUF

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes152downloads
Model Card

<p align="center"> <img src="https://raw.githubusercontent.com/Worth-Doing/brand-assets/main/png/variants/04-horizontal.png" alt="worthdoing" width="400"/> </p> <p align="center"><strong>Author: Simon-Pierre Boucher</strong></p>

<p align="center"> <img src="https://img.shields.io/badge/Format-GGUF-blue?style=for-the-badge" alt="GGUF"/> <img src="https://img.shields.io/badge/Params-7B-orange?style=for-the-badge" alt="Parameters"/> <img src="https://img.shields.io/badge/Platform-AppleSilicon-black?style=for-the-badge&logo=apple" alt="Apple Silicon"/> <img src="https://img.shields.io/badge/License-Apache2.0-green?style=for-the-badge" alt="License"/> <img src="https://img.shields.io/badge/Quantizedby-worthdoing-purple?style=for-the-badge" alt="worthdoing"/> </p> <p align="center"> <img src="https://img.shields.io/badge/Q4KM-3.7GB-brightgreen?style=flat-square" alt="Q4KM"/> <img src="https://img.shields.io/badge/Q5_KM-4.3GB-yellow?style=flat-square" alt="Q5KM"/> <img src="https://img.shields.io/badge/Q8_0-6.5GB-red?style=flat-square" alt="Q8_0"/> </p>

Qwen2.5-Coder-7B-Instruct - GGUF Quantized by worthdoing

Quantized for local Mac inference (Apple Silicon / Metal) by worthdoing

About

This is a GGUF quantized version of Qwen2.5-Coder-7B-Instruct, optimized for running locally on Apple Silicon Macs with llama.cpp, Ollama, or LM Studio.

Description

Qwen's dedicated coding model. Top-tier code generation and understanding.

Available Quantizations

FileQuantBPWSizeUse Case
qwen2.5-coder-7b-instruct-Q4_K_M-worthdoing.ggufQ4KM4.58~3.7 GBRecommended - Best quality/size ratio
qwen2.5-coder-7b-instruct-Q5_K_M-worthdoing.ggufQ5KM5.33~4.3 GBHigher quality, still fast
qwen2.5-coder-7b-instruct-Q8_0-worthdoing.ggufQ8_07.96~6.5 GBNear-original quality

How to Use

With Ollama

bash
# Create a Modelfile
cat > Modelfile <<'MODELEOF'
FROM ./qwen2.5-coder-7b-instruct-Q4_K_M-worthdoing.gguf
MODELEOF

ollama create qwen2.5-coder-7b-instruct -f Modelfile
ollama run qwen2.5-coder-7b-instruct

With llama.cpp

bash
llama-cli -m qwen2.5-coder-7b-instruct-Q4_K_M-worthdoing.gguf -p "Your prompt here" -ngl 99

With LM Studio

  1. 1.Download the GGUF file
  2. 2.Open LM Studio -> My Models -> Import
  3. 3.Select the GGUF file and start chatting

Quantization Method

Our quantization pipeline (corelm-model v1.0) follows a rigorous multi-step process to ensure maximum quality and compatibility:

Step 1 — Download & Validation

  • Model weights are downloaded from HuggingFace Hub in SafeTensors format (.safetensors)
  • Legacy formats (.bin, .pt) are excluded to ensure clean, verified weights
  • Tokenizer, configuration, and all metadata are preserved

Step 2 — Conversion to GGUF F16 Baseline

  • The original model is converted to GGUF format at FP16 precision using convert_hf_to_gguf.py from llama.cpp
  • This lossless baseline preserves the full original model quality
  • Architecture-specific tensors (attention, FFN, embeddings, MoE routing) are mapped to their GGUF equivalents

Step 3 — K-Quant Quantization

  • The F16 baseline is quantized using llama-quantize with k-quant methods
  • K-quants use a mixed-precision approach: more important layers (attention, output) retain higher precision, while less sensitive layers (FFN) are compressed more aggressively
  • Each quantization level offers a different quality/size tradeoff:
MethodBits per WeightStrategy
Q4_K_M~4.58 bpwMixed 4/5-bit. Attention & output layers use Q5K, FFN layers use Q4K. Best balance of quality and size.
Q5_K_M~5.33 bpwMixed 5/6-bit. Attention & output layers use Q6K, FFN layers use Q5K. Higher quality with moderate size increase.
Q8_0~7.96 bpwUniform 8-bit. All layers quantized to 8-bit. Near-lossless quality, largest file size.

Step 4 — Metadata Injection

  • Custom metadata is embedded directly in each GGUF file:
  • general.quantized_by: worthdoing
  • general.quantization_version: corelm-1.0
  • This ensures full traceability and provenance of every quantized file

Tools & Environment

  • llama.cpp: Used for both conversion and quantization — the industry-standard open-source LLM inference engine
  • Target platform: Apple Silicon Macs (M1/M2/M3/M4) with Metal GPU acceleration
  • Inference runtimes: Compatible with llama.cpp, Ollama, LM Studio, koboldcpp, and any GGUF-compatible runtime

Recommended Hardware

QuantMin RAMRecommended
Q4KM4 GBMac with 8 GB+ RAM
Q5KM5 GBMac with 8 GB+ RAM
Q8_08 GBMac with 12 GB+ RAM

Tags

coding, code-generation, code-review


Quantized with corelm-model pipeline by worthdoing on 2026-04-17