CoolFace
Modelpublic

vladimir94/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes646downloads
Model Card

Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF

GGUF conversion of YuYu1015/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4, itself an NVFP4 W4A4 quantization of huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated.

This repository is intended for llama.cpp on NVIDIA Blackwell systems such as DGX Spark / GB10. The text model is provided as a main GGUF plus a separate MTP draft GGUF for llama.cpp speculative decoding. An optional mmproj file is also included for image input.

This is an uncensored/abliterated model. Use it responsibly.

Files

FileSizePurpose
Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf20.95 GiBMain text model, MTP excluded
mtp-Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf3.48 GiBSeparate MTP draft model for llama.cpp
mmproj-Huihui-Qwen3.6-35B-A3B-abliterated-F16.gguf0.84 GiBOptional vision projector

The BF16 suffix means the converter used BF16 for non-NVFP4 tensors and fallback tensors. The source checkpoint is YuYu1015's NVFP4 packed model, and the llama.cpp runtime used for testing reported BLACKWELL_NATIVE_FP4 = 1.

Conversion Notes

Converted from the downloaded Hugging Face safetensors checkpoint with llama.cpp:

  • —llama.cpp commit: 65ef50a0a4bb240211a41d43c957ae6313af6841
  • —llama.cpp version output: version: 1 (65ef50a)
  • —Platform used for conversion/test: NVIDIA DGX Spark / GB10, Linux aarch64
  • —Runtime build reported: BLACKWELL_NATIVE_FP4 = 1

Main model and MTP were split deliberately:

  • —--no-mtp was used for the main model.
  • —--mtp was used on a second pass to create the standalone draft GGUF.
  • —This split duplicates some shared tensors, but lets llama.cpp load the MTP head as --model-draft.

Equivalent commands:

bash
SRC=/path/to/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4
OUT=/path/to/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-GGUF
LLAMA=/path/to/llama.cpp

python3 "$LLAMA/convert_hf_to_gguf.py" "$SRC" \
  --outfile "$OUT/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf" \
  --outtype bf16 \
  --no-mtp

python3 "$LLAMA/convert_hf_to_gguf.py" "$SRC" \
  --outfile "$OUT/mtp-Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf" \
  --outtype bf16 \
  --mtp

The vision projector needed a small conversion wrapper because this checkpoint stores visual tensors under model.language_model.visual.*, while the llama.cpp Qwen3-VL mmproj converter expects visual.*. The wrapper remaps only that prefix and then calls convert_hf_to_gguf.py.

Equivalent mmproj command after applying that prefix wrapper:

bash
python3 convert_qwen35_mmproj.py "$SRC" \
  --outfile "$OUT/mmproj-Huihui-Qwen3.6-35B-A3B-abliterated-F16.gguf" \
  --outtype f16 \
  --mmproj

Recommended llama.cpp Settings

Text-only, 262K context, MTP enabled:

bash
llama-server \
  --model Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf \
  --model-draft mtp-Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-BF16.gguf \
  --alias huihui-qwen3.6-35b-a3b-abliterated-uncensored-nvfp4-mtp \
  --ctx-size 262144 \
  --parallel 8 \
  --batch-size 8192 \
  --ubatch-size 2048 \
  --flash-attn on \
  --n-gpu-layers all \
  --kv-unified \
  --cont-batching \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --cache-type-k-draft q8_0 \
  --cache-type-v-draft q8_0 \
  --spec-type draft-mtp \
  --spec-draft-n-max 1 \
  --draft-p-min 0.3 \
  --cache-ram 8192 \
  --reasoning off

For vision, add:

bash
--mmproj mmproj-Huihui-Qwen3.6-35B-A3B-abliterated-F16.gguf

For the tested deployment, the non-vision preset used 8 slots. The vision preset used 2 slots, because loading mmproj adds memory and disables some cache-reuse behavior in llama.cpp multimodal mode.

MTP Parameter Testing

All numbers below are llama.cpp generation throughput on DGX Spark / GB10 at 262K context with the text-only GGUF, --flash-attn on, Q8_0 KV cache, and native Blackwell FP4 support enabled. Continuous batching was enabled for the live llama-server tests and was verified with MTP. These are local smoke/throughput tests, not a formal benchmark suite.

SetupThroughputNotes
No speculative decoding~31.0 tok/sBaseline
MTP, --spec-draft-n-max 1, --draft-p-min 0.0~34.76 tok/sAcceptance ~71.7%
MTP, --spec-draft-n-max 1, --draft-p-min 0.3~34.84 tok/sAcceptance ~78.3%; selected setting
MTP, --spec-draft-n-max 2, --draft-p-min 0.0~33.99 tok/sSlower than n=1
MTP, --spec-draft-n-max 2, --draft-p-min 0.3~33.07 tok/sSlower than n=1
MTP, --spec-draft-n-max 3, --draft-p-min 0.0~31.88 tok/sClose to baseline
MTP, --spec-draft-n-max 4, --draft-p-min 0.0~29.15 tok/sSlower than baseline

Recommended MTP setting for this GGUF:

bash
--spec-type draft-mtp --spec-draft-n-max 1 --draft-p-min 0.3

Higher MTP depths were not useful in the local tests. We stopped the sweep after n=4 because throughput was already declining, and kept n=1.

DFlash was not benchmarked for this GGUF/llama.cpp release. The source NVFP4 model card discusses DFlash for vLLM, but this repository is focused on llama.cpp with the model's built-in MTP draft.

DGX Spark / GB10 Throughput and Memory

Tested on a DGX Spark / GX10 with 128GB unified memory.

One-slot temporary llama.cpp test:

  • —Best observed setting: MTP n_max=1, draft_p_min=0.3
  • —Generation throughput: ~34.8 tok/s

Eight-slot live llama-server deployment:

  • —Server loaded with --parallel 8, --ctx-size 262144, --kv-unified, --cont-batching
  • —Each slot reported n_ctx = 262144
  • —Single active generation in the 8-slot continuous-batching server: ~30.6 to ~31.2 tok/s
  • —NVIDIA accounting for the worker process: ~30.4 GiB
  • —The separate llama-server controller process used about 170 MiB

The 8-slot throughput number is a single active request on an 8-slot server, not an 8-concurrent aggregate throughput benchmark.

Because --kv-unified is enabled, filling slots with context does not allocate eight separate full-size KV buffers. The live worker stayed around 30-31 GiB by nvidia-smi. Host memory can still grow due to prompt/checkpoint caching; the tested config set --cache-ram 8192.

Vision / mmproj Status

The optional mmproj file was converted and a local image-input smoke test succeeded. The test image had a red/blue split, and the model correctly described red and blue side-by-side.

Recommended usage:

  • —Use the text-only model for normal assistant/chat workloads.
  • —Load mmproj only when image input is needed.
  • —Treat vision support as available but less extensively benchmarked than text generation.

Known multimodal caveat from local llama.cpp logs:

  • —With mmproj loaded, llama.cpp reported that cache reuse is not supported by multimodal mode and disabled it.

Prompt Cache and Cache Reuse Notes

The tested llama.cpp deployment used:

bash
--cache-ram 8192
--kv-unified
--cont-batching

Prompt/checkpoint caching worked in the sense that idle slots were saved and restored from the prompt cache. However, for this hybrid Qwen3.6/GDN context, llama.cpp also logged that cache_reuse was not supported and disabled cache reuse. In practice, do not expect every repeated prompt to show classic prefix-cache hit behavior.

Safety

This is an abliterated/uncensored model. It may produce unsafe, offensive, or policy-violating content. Users are responsible for deployment choices, filtering, logging, and compliance with applicable law.

Credits