CoolFace
Modelpublic

michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF

sourceHugging Faceupdated 4mo agoView on Hugging Face
12likes366kdownloads
Model Card

Qwen3.6-35B-A3B-NVFP4-MTP-GGUF

This repo contains two experimental NVFP4 GGUF quantizations of Qwen3.6-35B-A3B for llama.cpp.<BR> This was quantized using my experimental <A HREF="https://github.com/michaelw9999/advanced-gguf-quantizer/">advanced-gguf-quantizer</A> tool.<BR> Both models were imatrix calibrated for the first time using a new custom dataset that I am evaluating.

This repository contains two NVFP4 variants:

VariantFileBest forNotes
TURBO`Qwen3.6-35B-A3B-NVFP4-MTP-TURBO.gguf`Max speedMore NVFP4. Lower quality metrics.
HQ`Qwen3.6-35B-A3B-NVFP4-MTP-HQ.gguf`Better qualityMore tensors promoted. Slightly slower.

Quality & Speed Results

All PPL/KLD results were measured against the same BF16 wikitest KLD base, and then compared to the official NVFP4 release by NVIDIA.

MetricTURBOHQNVIDIA-NVFP4
Size18.56 GiB18.64 GiB22.20 GiB
Mean PPL(Q)6.9873926.8977967.014030
Mean PPL(Q)-PPL(base)0.2685510.178955—
Mean PPL ratio1.0399701.0266351.043935
Mean ln(PPL ratio)0.0391920.026286—
Mean KLD0.0632280.0507590.066331
99.9% KLD1.9241471.5651431.560988
99.0% KLD0.5985190.4883870.495896
95.0% KLD0.2210300.1788890.207580
Max KLD11.94657110.0939116.972712
Same top p89.023%90.255%87.608%
Top flip weight0.0120680.009575—
pp51211593.57 t/s10936.20 t/s10426.32 t/s
tg128271.21 t/s270.49 t/s221.86 t/s

Evaluation Results

Further evaluation tests are underway to identify real world performance differences between TURBO and HQ.

BenchmarkSamplesTURBOHQNVIDIA-NVFP4
GSM8K10398%98%97%
HellaSwag10089%89%89%
HumanEval16496.34%95.12%95.12%