CoolFace
Modelpublic

lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-IQ4_XS-GGUF

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
32likes9.5kdownloads
Model Card

Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-IQ4_XS-GGUF

GGUF quantizations of `lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled` for use with llama.cpp and LM Studio.

The base model is a reasoning-distilled variant of Qwen3.6-35B-A3B fine-tuned to imitate the chain-of-thought style of Claude Opus 4.7. It thinks in explicit <think>...</think> blocks before producing the final answer.

Quant files

See the file list for all available quant levels. Common choices:

FileQuantApprox sizeUse case
*.IQ4_XS.ggufIQ4_XS~18 GBSmallest quant with good quality — default pick for LM Studio
*.Q4_K_M.ggufQ4KM~21 GBBalanced quality / size
*.Q5_K_M.ggufQ5KM~25 GBHigher quality
*.Q8_0.ggufQ8_0~35 GBNear-lossless

Running in llama.cpp

bash
llama-server \
  -m Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled.IQ4_XS.gguf \
  --host 127.0.0.1 --port 18081 \
  -c 32768 -fa on \
  --cache-type-k q8_0 --cache-type-v turbo4

Running in LM Studio

Search for lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-IQ4_XS-GGUF inside LM Studio's model browser and pick the quant that fits your RAM/VRAM. The model should appear automatically once HF indexes this repo.

License

Apache 2.0, inherited from the base model. See `lordx64/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled` for training details, evaluations, and intended use.