vladimir94/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF
Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF
GGUF conversion of YuYu1015/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4, itself an NVFP4 W4A4 quantization of huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated.
This repository is intended for llama.cpp on NVIDIA Blackwell systems such as DGX Spark / GB10. The text model is provided as a main GGUF plus a separate MTP draft GGUF for llama.cpp speculative decoding. An optional mmproj file is also included for image input.
This is an uncensored/abliterated model. Use it responsibly.
Files
The BF16 suffix means the converter used BF16 for non-NVFP4 tensors and fallback tensors. The source checkpoint is YuYu1015's NVFP4 packed model, and the llama.cpp runtime used for testing reported BLACKWELL_NATIVE_FP4 = 1.
Conversion Notes
Converted from the downloaded Hugging Face safetensors checkpoint with llama.cpp:
- llama.cpp commit:
65ef50a0a4bb240211a41d43c957ae6313af6841 - llama.cpp version output:
version: 1 (65ef50a) - Platform used for conversion/test: NVIDIA DGX Spark / GB10, Linux aarch64
- Runtime build reported:
BLACKWELL_NATIVE_FP4 = 1
Main model and MTP were split deliberately:
--no-mtpwas used for the main model.--mtpwas used on a second pass to create the standalone draft GGUF.- This split duplicates some shared tensors, but lets llama.cpp load the MTP head as
--model-draft.
Equivalent commands:
SRC=/path/to/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4
OUT=/path/to/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF
LLAMA=/path/to/llama.cpp
python3 "$LLAMA/convert_hf_to_gguf.py" "$SRC" \
--outfile "$OUT/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf" \
--outtype bf16 \
--no-mtp
python3 "$LLAMA/convert_hf_to_gguf.py" "$SRC" \
--outfile "$OUT/mtp-Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf" \
--outtype bf16 \
--mtpThe vision projector needed a small conversion wrapper because this checkpoint stores visual tensors under model.language_model.visual.*, while the llama.cpp Qwen3-VL mmproj converter expects visual.*. The wrapper remaps only that prefix and then calls convert_hf_to_gguf.py.
Equivalent mmproj command after applying that prefix wrapper:
python3 convert_qwen35_mmproj.py "$SRC" \
--outfile "$OUT/mmproj-Huihui-Qwen3.6-35B-A3B-abliterated-F16.gguf" \
--outtype f16 \
--mmprojRecommended llama.cpp Settings
Text-only, 262K context, MTP enabled:
llama-server \
--model Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf \
--model-draft mtp-Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf \
--alias huihui-qwen3.6-35b-a3b-abliterated-uncensored-nvfp4-mtp \
--ctx-size 262144 \
--parallel 8 \
--batch-size 8192 \
--ubatch-size 2048 \
--flash-attn on \
--n-gpu-layers all \
--kv-unified \
--cont-batching \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--cache-type-k-draft q8_0 \
--cache-type-v-draft q8_0 \
--spec-type draft-mtp \
--spec-draft-n-max 1 \
--draft-p-min 0.3 \
--cache-ram 8192 \
--reasoning offFor vision, add:
--mmproj mmproj-Huihui-Qwen3.6-35B-A3B-abliterated-F16.ggufFor the tested deployment, the non-vision preset used 8 slots. The vision preset used 2 slots, because loading mmproj adds memory and disables some cache-reuse behavior in llama.cpp multimodal mode.
MTP Parameter Testing
All numbers below are llama.cpp generation throughput on DGX Spark / GB10 at 262K context with the text-only GGUF, --flash-attn on, Q8_0 KV cache, and native Blackwell FP4 support enabled. Continuous batching was enabled for the live llama-server tests and was verified with MTP. These are local smoke/throughput tests, not a formal benchmark suite.
Recommended MTP setting for this GGUF:
--spec-type draft-mtp --spec-draft-n-max 1 --draft-p-min 0.3Higher MTP depths were not useful in the local tests. We stopped the sweep after n=4 because throughput was already declining, and kept n=1.
DFlash was not benchmarked for this GGUF/llama.cpp release. The source NVFP4 model card discusses DFlash for vLLM, but this repository is focused on llama.cpp with the model's built-in MTP draft.
DGX Spark / GB10 Throughput and Memory
Tested on a DGX Spark / GX10 with 128GB unified memory.
One-slot temporary llama.cpp test:
- Best observed setting: MTP
n_max=1,draft_p_min=0.3 - Generation throughput: ~34.8 tok/s
Eight-slot live llama-server deployment:
- Server loaded with
--parallel 8,--ctx-size 262144,--kv-unified,--cont-batching - Each slot reported
n_ctx = 262144 - Single active generation in the 8-slot continuous-batching server: ~30.6 to ~31.2 tok/s
- NVIDIA accounting for the worker process: ~30.4 GiB
- The separate llama-server controller process used about 170 MiB
The 8-slot throughput number is a single active request on an 8-slot server, not an 8-concurrent aggregate throughput benchmark.
Because --kv-unified is enabled, filling slots with context does not allocate eight separate full-size KV buffers. The live worker stayed around 30-31 GiB by nvidia-smi. Host memory can still grow due to prompt/checkpoint caching; the tested config set --cache-ram 8192.
Vision / mmproj Status
The optional mmproj file was converted and a local image-input smoke test succeeded. The test image had a red/blue split, and the model correctly described red and blue side-by-side.
Recommended usage:
- Use the text-only model for normal assistant/chat workloads.
- Load
mmprojonly when image input is needed. - Treat vision support as available but less extensively benchmarked than text generation.
Known multimodal caveat from local llama.cpp logs:
- With
mmprojloaded, llama.cpp reported that cache reuse is not supported by multimodal mode and disabled it.
Prompt Cache and Cache Reuse Notes
The tested llama.cpp deployment used:
--cache-ram 8192
--kv-unified
--cont-batchingPrompt/checkpoint caching worked in the sense that idle slots were saved and restored from the prompt cache. However, for this hybrid Qwen3.6/GDN context, llama.cpp also logged that cache_reuse was not supported and disabled cache reuse. In practice, do not expect every repeated prompt to show classic prefix-cache hit behavior.
Safety
This is an abliterated/uncensored model. It may produce unsafe, offensive, or policy-violating content. Users are responsible for deployment choices, filtering, logging, and compliance with applicable law.
Credits
- Original base: Qwen/Qwen3.6-35B-A3B
- Abliterated model: huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
- Source NVFP4 checkpoint: YuYu1015/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4
- GGUF conversion and DGX Spark llama.cpp testing: vladimir94
