jan1k/Qwen3.6-35B-A3B-Uncensored-Genesis-NVFP4-GGUF
Qwen3.6-35B-A3B-Uncensored-Genesis — NVFP4 GGUF
NVFP4 quantisation of LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Final-GGUF, built with advanced-gguf-quantizer (a llama.cpp fork focused on NVFP4/MXFP6 quantization).
This is the non-Hermes version — Genesis tensor repair on the HauhauCS uncensored base, without the Hermes finetune transfer. For the Hermes version, see jan1k/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-Final-NVFP4-GGUF.
Files
v4 — Recommended (inline scales, LM Studio compatible)
Re-quantised 2026-09-24 from Luffy's updated Q8KP source.
v4 uses --nvfp4-inline-scales-only to emit native inline UE4M3 scales without separate .scale/.input_scale tensors. This is required for LM Studio / Pelican and other runtimes that do not support the extended NVFP4 scale tensor contract.
Imatrix variants will follow after the plain v4 is confirmed working in target runtimes.
Source
Two-step pipeline: Q8KP -> F16 intermediate -> NVFP4. The F16 step gives the NVFP4 encoder clean data (no Q8_0 quantization noise).
Tensor mix (v4)
Format fix (v3/v2 -> v4)
The older v2/v3 files use the extended NVFP4 format with separate .scale and .input_scale tensors. Runtimes like LM Studio / Pelican only read inline UE4M3 scales and ignore the separate scale tensors, producing garbage output.
v4 uses --nvfp4-inline-scales-only to:
- Skip writing
.scaleand.input_scaletensors - Force
tensor_scale=1.0(no pre-scaling) - Force
input_scale=identity - Make inline UE4M3 scales self-contained
Tensor protection policy
F16 singular-collapse protection:
F32 architecture-specific protection:
Forced NVFP4 (do not push lower):
Usage
llama-cli -m Qwen3.6-35B-A3B-Uncensored-Genesis-NVFP4-v4.gguf \
--mmproj mmproj-Qwen3.6-35B-A3B-Uncensored-Genesis-Final-F16.gguf \
--jinja -c 131072 -ngl 99- Set K cache and V cache quantization to F16
- Set GPU offload to maximum, active experts to 8
- Set number of layers for which to force MoE weights onto CPU to 40
Hardware
- Blackwell (RTX 50xx): native FP4 path, fastest
- Ampere (RTX 30xx): NVFP4 inference works via fallback kernels
- Quantisation was done CPU-only (Ampere CUDA NVFP4 encoder hangs on MoE)
Reproducibility
# Step 1: Q8_K_P -> F16 intermediate
llama-quantize --allow-requantize \
Qwen3.6-35B-A3B-Uncensored-Genesis-Q8_K_P.gguf temp_f16.gguf F16 6
# Step 2: F16 -> NVFP4 v4 (inline scales only, CPU-only)
CUDA_VISIBLE_DEVICES="" llama-quantize \
--allow-requantize \
--nvfp4-inline-scales-only \
--tensor-type-file tensor_types_protection.txt \
temp_f16.gguf \
Qwen3.6-35B-A3B-Uncensored-Genesis-NVFP4-v4.gguf \
NVFP4 6
# Cleanup
rm temp_f16.ggufCredits
- Base model: HauhauCS (0/465 refusals)
- Genesis tensor repair: LuffyTheFox
- NVFP4 quantisation: jan1k
- Quantiser: advanced-gguf-quantizer
