CoolFace
Modelpublic

mlx-community/Qwen3.6-27B-F451-AND-TRI-Polar-Ultra-Pro-Writer-Uncensored-Heretic-OptiQ-4bit

sourceHugging Faceapache-2.0updated 12d agoView on Hugging Face
2likes1.1kdownloads
Model Card

Qwen3.6-27B-F451-AND-TRI-Polar-Ultra-Pro-Writer-Uncensored-Heretic-OptiQ-4bit

An OptiQ mixed-precision quantization of DavidAU/Qwen3.6-27B-F451-AND-TRI-Polar-Ultra-Pro-Writer-Uncensored-Heretic, built for Apple Silicon with MLX.

OptiQ assigns a per-layer bit-width (4-bit or 8-bit here) instead of quantizing every layer the same, so the layers that need the precision keep it.

  • —Method: OptiQ data-driven mixed precision (optiq_mixed_precision)
  • —Average: ~5.6 bits per weight (220 layers at 8-bit, 276 at 4-bit)
  • —Architecture: qwen3_5 (Qwen3.6-27B)

The per-layer allocation is reused from mlx-community/Qwen3.6-27B-OptiQ-4bit, which shares the same Qwen3.6-27B architecture, so no separate sensitivity pass was run for this merge. This is a text-only quant of the language tower. The tokenizer is copied from the base model unchanged.

Use

bash
pip install mlx-optiq
optiq serve --model mlx-community/Qwen3.6-27B-F451-AND-TRI-Polar-Ultra-Pro-Writer-Uncensored-Heretic-OptiQ-4bit

Or load it directly with mlx_lm:

python
from mlx_lm import load, generate
model, tokenizer = load("mlx-community/Qwen3.6-27B-F451-AND-TRI-Polar-Ultra-Pro-Writer-Uncensored-Heretic-OptiQ-4bit")

See mlx-optiq.com for the CLI, the local Lab UI, and the OptiQ Code agent. The base model and its behavior are DavidAU's; this repo only changes the quantization.