michaelw9999/Qwen3.6-35B-A3B-NVFP4-MTP-GGUF
12366k
Qwen3.6-35B-A3B-NVFP4-MTP-GGUF
This repo contains two experimental NVFP4 GGUF quantizations of Qwen3.6-35B-A3B for llama.cpp.<BR> This was quantized using my experimental <A HREF="https://github.com/michaelw9999/advanced-gguf-quantizer/">advanced-gguf-quantizer</A> tool.<BR> Both models were imatrix calibrated for the first time using a new custom dataset that I am evaluating.
This repository contains two NVFP4 variants:
Quality & Speed Results
All PPL/KLD results were measured against the same BF16 wikitest KLD base, and then compared to the official NVFP4 release by NVIDIA.
Evaluation Results
Further evaluation tests are underway to identify real world performance differences between TURBO and HQ.
