CoolFace
Modelpublic

magiccodingman/Qwen3-4B-Instruct-2507-Unsloth-MagicQuant-Hybrid-GGUF

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
8likes364downloads
Model Card

MagicQuant GGUF Hybrids - Qwen3 4B Instruct 2507

(DEPRECIATED - Part of MagicQuant v1.0 which had significant flaws. Please utilize v2.0 which is production ready)

MagicQuant is an automated quantization, benchmarking, and evolutionary hybrid-GGUF search system for LLMs.

Each release includes models optimized to outperform standard baseline quants (Q8, Q6, Q5, Q4). If a baseline GGUF exists in this repo, the evolutionary engine couldn’t beat it. If a baseline is missing, it’s because a hybrid configuration outperformed it so completely that including the baseline would've been pointless.

These hybrid GGUFs are built to be as small, fast, and low-drift as possible while preserving model capability.

To dive deeper into how MagicQuant works, see the main repo: MagicQuant on GitHub (by MagicCodingMan)

Notes:

  • —The HuggingFace hardware compatibility where it shows the bits is usually wrong. It doesn't understand hybrid mixes, so don't trust it.
  • —Naming scheme can be found on the MagicQuant Wiki.
  • —(tips) Less precision loss means less brain damage. More TPS means faster! Smaller is always better right?

Precision Loss Guide

  • —0–0.1% → God-tier, scientifically exact
  • —0.1–1% → True near-lossless, agent-ready
  • —1–3% → Minimal loss, great for personal use
  • —3–5% → Borderline, but still functional
  • —5%+ → Toys, not tools, outside MagicQuant’s scope

Learn more about precision loss here.

Table - File Size + TPS + Avg Precision Loss

model_namefile_size_gbbench_tpsavg_prec_loss
mxfp4_moe-K-B16-QO-Q6K-EUD-Q8_03.98373.480.0533%
mxfp4_moe-O-Q5K-EQKUD-Q6K3.03428.370.1631%
mxfp4_moe-QUD-IQ4NL-KO-MXFP4-E-Q8_02.28411.490.7356%
mxfp4_moe-K-B16-QU-IQ4NL-O-MXFP4-E-Q5K-D-Q6K2.62467.790.8322%
IQ4_NL2.23426.860.8996%
mxfp4_moe-EQUD-IQ4NL-KO-MXFP42.10518.152.0904%

Table - PPL Columns

model_namegengen_ercodecode_ermathmath_er
mxfp4_moe-K-B16-QO-Q6K-EUD-Q8_08.87660.20531.54630.01226.71190.1368
mxfp4_moe-O-Q5K-EQKUD-Q6K8.85640.20361.54730.01226.69760.1358
mxfp4_moe-QUD-IQ4NL-KO-MXFP4-E-Q8_09.01270.20571.55460.01196.69190.1331
mxfp4_moe-K-B16-QU-IQ4NL-O-MXFP4-E-Q5K-D-Q6K9.04900.20961.55350.01216.72210.1358
IQ4_NL8.99480.20721.56000.01236.74840.1362
mxfp4_moe-EQUD-IQ4NL-KO-MXFP49.21040.21061.55980.01196.82610.1363

Table - Precision Loss Columns

model_nameloss_generalloss_codeloss_math
mxfp4_moe-K-B16-QO-Q6K-EUD-Q8_00.07200.03880.0492
mxfp4_moe-O-Q5K-EQKUD-Q6K0.29940.02590.1640
mxfp4_moe-QUD-IQ4NL-KO-MXFP4-E-Q8_01.46010.49780.2489
mxfp4_moe-K-B16-QU-IQ4NL-O-MXFP4-E-Q5K-D-Q6K1.86870.42670.2012
IQ4_NL1.25860.84690.5933
mxfp4_moe-EQUD-IQ4NL-KO-MXFP43.68570.83391.7515

Baseline Models (Reference)

Table - File Size + TPS + Avg Precision Loss

model_namefile_size_gbbench_tpsavg_prec_loss
BF167.50254.700.0000%
Q8_03.99362.480.0724%
Q6_K3.08397.920.2492%
Q5_K2.69385.170.7920%
IQ4_NL2.23426.860.8996%
Q4KM2.33377.190.9376%
MXFP4_MOE2.00467.138.2231%

Table - PPL Columns

model_namegengen_ercodecode_ermathmath_er
BF168.88300.20561.54690.01226.70860.1369
Q8_08.87540.20531.54880.01236.70800.1367
Q6_K8.84410.20341.54520.01216.69520.1357
Q5_K8.97070.20791.55420.01236.77010.1384
IQ4_NL8.99480.20721.56000.01236.74840.1362
Q4KM8.94460.20511.56940.01256.75320.1371
MXFP4_MOE9.87990.22821.61220.01307.32750.1494

Table - Precision Loss Columns

model_nameloss_generalloss_codeloss_math
BF160.00000.00000.0000
Q8_00.08560.12280.0089
Q6_K0.43790.10990.1997
Q5_K0.98730.47190.9167
IQ4_NL1.25860.84690.5933
Q4KM0.69351.45450.6648
MXFP4_MOE11.22264.22139.2255

Support

I’m a solo developer working full time for myself to achieve my dream, pouring nights and weekends into open protocols and tools that I hope make the world a little better. If you chip in, you're helping me keep the lights on while I keep shipping.

Click here to see ways to support - BTC, Paypal, GitHub sponsors.

Or, just drop a like on the repo :)