cesarsal1nas/Huihui-Qwen3.5-35B-A3B-abliterated-Q4_K_M-GGUF
2315
[!WARNING] A significantly improved version of this model is available. This repo quantizes huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated — the standard abliteration. A newer fine-tune of the same architecture, trained in the style of Claude 4.6 Opus, has since been released and produces noticeably richer, more expressive outputs. ➡️ Recommended upgrade: [Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated-Q4_K_M-GGUF](https://huggingface.co/cesarsal1nas/Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated-Q4_K_M-GGUF) Same architecture, same Q4KM quantization, same VRAM footprint — just a better fine-tune. This repo will remain available for reference.
Huihui-Qwen3.5-35B-A3B-abliterated — Q4KM GGUF
This is a Q4KM GGUF quantization of huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated.
Refer to the original model card for full details, usage warnings, and licensing information.
Details
Usage with llama.cpp
llama-cli \
--hf-repo cesarsal1nas/Huihui-Qwen3.5-35B-A3B-abliterated-Q4_K_M-GGUF \
--hf-file huihui-qwen3.5-35b-a3b-abliterated-Q4_K_M.gguf \
-p "Tell me about the universe"llama-server \
--hf-repo cesarsal1nas/Huihui-Qwen3.5-35B-A3B-abliterated-Q4_K_M-GGUF \
--hf-file huihui-qwen3.5-35b-a3b-abliterated-Q4_K_M.gguf \
-c 8192Usage with Ollama
Requires Ollama with qwen35moe support. See PR #14506 for the architecture patch.
ollama run hf.co/cesarsal1nas/Huihui-Qwen3.5-35B-A3B-abliterated-Q4_K_M-GGUFCredits
- Abliteration by huihui-ai — Huihui-Qwen3.5-35B-A3B-abliterated
- Base model by Qwen — Qwen3.5-35B-A3B
- Quantization by cesarsal1nas
