him0413/Qwen3.8-Flash-Next-Uncensored-Q4_K_M-MTP
Qwen3.8-Flash-Next-Uncensored Q4KM (Integrated MTP)
An integrated MTP (Multi-Token Prediction) GGUF build of the abliterated Qwen3.8-Flash-Next-Uncensored model, quantized to Q4_K_M with the MTP draft head merged into the same 4-shard split — no sidecar file needed.
Model Details
Usage (llama.cpp with qwen4exp support)
./build-vulkan/bin/llama-server \
--model Qwen3.8-Flash-Next-Uncensored-Q4_K_M-MTP-00001-of-00004.gguf \
--flash-attn on \
--spec-type draft-mtp \
--spec-draft-adaptive \
--spec-draft-n-min 0 \
--spec-draft-n-max 7 \
--spec-draft-p-min 0.75What's inside
This repo contains the MTP head tensors fused into the main GGUF:
blk.48.nextn.*— MTP embedding/hidden projections + normsblk.48.nextn.hc_head_*— draft head output mixer- Full 48-layer trunk with
nextn_predict_layers = 1
⚠️ Disclaimer
This model is an abliterated (refusal-removed) build. It will comply with harmful, unethical, or illegal requests the original Qwen3.8-Flash-Next would refuse. Released strictly for legitimate research — interpretability, AI-safety / refusal-mechanism study, red-teaming, and robustness evaluation. You assume full responsibility for how you use it and everything it generates; add your own safety and moderation layers before any deployment.
License
Qwen Community License 1.0 (see LICENSE), inherited from the base model `Qwen/Qwen3.8-Flash-Next`. Abliteration and quantization do not change the underlying license obligations.
Note: per the Qwen Community License, if you operate a Model as a Service or AI Work Assistant business commercially, you must obtain a separate license from Qwen before using this model or its derivatives for commercial purposes. See LICENSE for full terms.
