CoolFace
Modelpublic

worthdoing/Mixtral-8x7B-Instruct-v0.1-GGUF

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes39downloads
Model Card

<p align="center"> <img src="https://raw.githubusercontent.com/Worth-Doing/brand-assets/main/png/variants/04-horizontal.png" alt="worthdoing" width="400"/> </p> <p align="center"><strong>Author: Simon-Pierre Boucher</strong></p>

<p align="center"> <img src="https://img.shields.io/badge/Format-GGUF-blue?style=for-the-badge" alt="GGUF"/> <img src="https://img.shields.io/badge/Params-46.7B-orange?style=for-the-badge" alt="Parameters"/> <img src="https://img.shields.io/badge/Platform-AppleSilicon-black?style=for-the-badge&logo=apple" alt="Apple Silicon"/> <img src="https://img.shields.io/badge/License-Apache2.0-green?style=for-the-badge" alt="License"/> <img src="https://img.shields.io/badge/Quantizedby-worthdoing-purple?style=for-the-badge" alt="worthdoing"/> </p> <p align="center"> <img src="https://img.shields.io/badge/Q4KM-24.9GB-brightgreen?style=flat-square" alt="Q4KM"/> <img src="https://img.shields.io/badge/Q5_KM-29.0GB-yellow?style=flat-square" alt="Q5KM"/> <img src="https://img.shields.io/badge/Q8_0-43.3GB-red?style=flat-square" alt="Q8_0"/> </p>

Mixtral-8x7B-Instruct-v0.1 - GGUF Quantized by worthdoing

Quantized for local Mac inference (Apple Silicon / Metal) by worthdoing

About

This is a GGUF quantized version of Mixtral-8x7B-Instruct-v0.1, optimized for running locally on Apple Silicon Macs with llama.cpp, Ollama, or LM Studio.

Description

Mistral's MoE model. 46.7B total, 12B active. Fast with great quality.

Available Quantizations

FileQuantBPWSizeUse Case
mixtral-8x7b-instruct-v0.1-Q4_K_M-worthdoing.ggufQ4KM4.58~24.9 GBRecommended - Best quality/size ratio
mixtral-8x7b-instruct-v0.1-Q5_K_M-worthdoing.ggufQ5KM5.33~29.0 GBHigher quality, still fast
mixtral-8x7b-instruct-v0.1-Q8_0-worthdoing.ggufQ8_07.96~43.3 GBNear-original quality

How to Use

With Ollama

bash
# Create a Modelfile
cat > Modelfile <<'MODELEOF'
FROM ./mixtral-8x7b-instruct-v0.1-Q4_K_M-worthdoing.gguf
MODELEOF

ollama create mixtral-8x7b-instruct-v0.1 -f Modelfile
ollama run mixtral-8x7b-instruct-v0.1

With llama.cpp

bash
llama-cli -m mixtral-8x7b-instruct-v0.1-Q4_K_M-worthdoing.gguf -p "Your prompt here" -ngl 99

With LM Studio

  1. 1.Download the GGUF file
  2. 2.Open LM Studio -> My Models -> Import
  3. 3.Select the GGUF file and start chatting

Quantization Method

Our quantization pipeline (corelm-model v1.0) follows a rigorous multi-step process to ensure maximum quality and compatibility:

Step 1 — Download & Validation

  • —Model weights are downloaded from HuggingFace Hub in SafeTensors format (.safetensors)
  • —Legacy formats (.bin, .pt) are excluded to ensure clean, verified weights
  • —Tokenizer, configuration, and all metadata are preserved

Step 2 — Conversion to GGUF F16 Baseline

  • —The original model is converted to GGUF format at FP16 precision using convert_hf_to_gguf.py from llama.cpp
  • —This lossless baseline preserves the full original model quality
  • —Architecture-specific tensors (attention, FFN, embeddings, MoE routing) are mapped to their GGUF equivalents

Step 3 — K-Quant Quantization

  • —The F16 baseline is quantized using llama-quantize with k-quant methods
  • —K-quants use a mixed-precision approach: more important layers (attention, output) retain higher precision, while less sensitive layers (FFN) are compressed more aggressively
  • —Each quantization level offers a different quality/size tradeoff:
MethodBits per WeightStrategy
Q4_K_M~4.58 bpwMixed 4/5-bit. Attention & output layers use Q5K, FFN layers use Q4K. Best balance of quality and size.
Q5_K_M~5.33 bpwMixed 5/6-bit. Attention & output layers use Q6K, FFN layers use Q5K. Higher quality with moderate size increase.
Q8_0~7.96 bpwUniform 8-bit. All layers quantized to 8-bit. Near-lossless quality, largest file size.

Step 4 — Metadata Injection

  • —Custom metadata is embedded directly in each GGUF file:
  • —general.quantized_by: worthdoing
  • —general.quantization_version: corelm-1.0
  • —This ensures full traceability and provenance of every quantized file

Tools & Environment

  • —llama.cpp: Used for both conversion and quantization — the industry-standard open-source LLM inference engine
  • —Target platform: Apple Silicon Macs (M1/M2/M3/M4) with Metal GPU acceleration
  • —Inference runtimes: Compatible with llama.cpp, Ollama, LM Studio, koboldcpp, and any GGUF-compatible runtime

Recommended Hardware

QuantMin RAMRecommended
Q4KM32 GBMac with 49 GB+ RAM
Q5KM37 GBMac with 57 GB+ RAM
Q8_056 GBMac with 86 GB+ RAM

Tags

general, coding, moe, fast


Quantized with corelm-model pipeline by worthdoing on 2026-04-17