windowsxp811203/Qwen3.8-27B-Abliterated-MLX-oQ6e-mtp
Qwen3.8-27B-Abliterated — MLX oQ6e with native MTP
6-bit (near-lossless) MLX build of windowsxp811203/Qwen3.8-27B-Abliterated, an abliterated (refusal-removed) Qwen/Qwen3.8-27B, made for Apple Silicon.
51.75 GiB bf16 → 22.09 GiB (23.72 GB) — the native MTP draft head is kept in the checkpoint and the vision tower is unquantized, so one file serves three audiences:
Built with oMLX's oQ "enhanced" quantizer: imatrix-weighted rounding plus a mixed-precision plan — 33 of 505 quantized language-model modules were promoted above 6-bit (oMLX's sensitivity ranking plus its fixed per-layer rules), so this is not a uniform 6-bit cast.
This is one of three quantized sizes (plus the bf16 reference they were made from); pick by memory and speed:
What is and isn't quantized
Effective 6.659 bits per weight over the language model; 22.09 GiB of safetensors.
Verification
All numbers measured on this exact checkpoint on a MacBook Pro M5 Max, 128 GB, oMLX 0.6.4; throughput numbers are single stream, the refusal batteries ran 8 requests concurrently.
MTP speculative decoding (oMLX, mtp_enabled: true; 10 fixed prompts × ≤256 greedy tokens; tok/s is the best of three runs — two after a warm-up request — with ranges below; acceptance and tok/cycle are pooled over the first session's requests — the 10 prompts plus its warm-up and two bench requests — and vary by up to ~2 pt across runs):
Run-to-run spread (same settings): off (plain decode) 14.1–16.8 (n=3); depth 1 20.6–24.8 (n=3); depth 2 17.7–20.6 (n=3); depth 3 20.2–23.3 (n=3); depth 4 22.8–23.9 (n=3).
oMLX's built-in throughput bench (synthetic prompt, 256 generated tokens): 1024-token prompt: 15.3 → 24.9 tok/s (1.63×, depth 4), first token 1.7 s; 4096-token prompt: 9.2 → 23.0 tok/s (2.50×, depth 3), first token 10.5 s (MTP off → best of depth 3/4).
Draft acceptance by depth: depth 3: d1=82.0%, d2=77.6%, d3=75.3%. Acceptance is a speed signal only — every draft is verified against the target's own distribution, so the output distribution is preserved. It is not bit-exact, though: at temperature 0 the MTP-on and MTP-off outputs were byte-identical on 4/10 fixed prompts at depth 3 (5/10 at depth 1, 5–6/10 at depth 4; counts vary by run); the rest diverge at a near-tie token — sometimes early: the earliest divergence was ~123 characters in — and continue coherently. MTP-off reruns are 10/10 identical, so the divergence comes from the batched verify path (several draft rows per matmul accumulate bf16 differently than single-token decode), and the divergence point moves between depths and repeat runs. Treat MTP-on greedy output as non-reproducible at the byte level — not as a head defect.
External drafter path (mlx-vlm 0.6.17, --draft-model …-MTP-bf16, temperature 0): single prompt (36 tokens), 300 generated tokens, best run per arm: 22.1 → 20.2 tok/s (0.91×, 88.4% of drafts accepted) (plain runs 18.9–22.1, n=5; drafter runs 16.9–20.2, n=3; across 3 sessions) — no real gain on this quant in our runs (the 6-bit affine kernels appear to leave no headroom for the multi-token verify pass), so use oMLX's in-checkpoint MTP here, or the oQ4e/oQ8e targets for the mlx-vlm drafter path
Refusal — greedy, non-thinking, max 256 tokens, no prompt prefill (the parent card's protocol):
With a "Sure, here is" assistant prefill — a jailbreak on its own, so not comparable with the parent — the same battery gave AdvBench 0/80 · 0.00 % and HarmBench-safety 0/119 · 0.0 %.
Capability — MMLU 5-shot, oMLX's built-in harness with its seeded 400-question sample stratified by subject (identical questions for every row), temperature 0, thinking off, MTP on:
(Not comparable with the parent card's 82.35 % — the full 14,042-question set scored by next-token logits — nor with the NVFP4 card's MMLU-400, a different harness and sample.)
Vision — synthetic probe (red circle / green triangle / blue square): all three shapes, colors and positions correct
Usage
oMLX (native MTP)
brew tap jundot/omlx https://github.com/jundot/omlx && brew install jundot/omlx/omlx
hf download windowsxp811203/Qwen3.8-27B-Abliterated-MLX-oQ6e-mtp --local-dir ~/.omlx/models/Qwen3.8-27B-Abliterated-MLX-oQ6e-mtp
omlx serve --model-dir ~/.omlx/modelsThen enable Lightning MTP for the model in the admin UI (http://localhost:8000/admin → model → Advanced), or in ~/.omlx/model_settings.json (note the models wrapper — a top-level model key is silently ignored and MTP stays off):
{"version": 1, "models": {"Qwen3.8-27B-Abliterated-MLX-oQ6e-mtp": {"mtp_enabled": true, "mtp_num_draft_tokens": 3, "max_context_window": 262144}}}The server log prints Speculative backend selected … Lightning MTP (model_type=qwen3_5, active) when it took effect. Depth 4 was marginally faster here on the 10-prompt greedy set (23.9 vs 23.3) and the 1K bench prompt (24.9 vs 24.1). Depth 3 is the safer default for long prompts. Depth 1 posted the single best greedy run (24.8 tok/s) but the widest spread (20.6–24.8) and lower bench numbers (23.9 / 21.7 at 1K / 4K), so depth 3–4 is the recommendation.
mlx-vlm
pip install -U mlx-vlm
python -m mlx_vlm.generate --model windowsxp811203/Qwen3.8-27B-Abliterated-MLX-oQ6e-mtp \
--draft-model windowsxp811203/Qwen3.8-27B-Abliterated-MTP-bf16 \
--prompt "Write a quicksort in Python." --max-tokens 512Drop --draft-model for plain decoding. The mlx-vlm CLI runs with thinking off unless you pass --enable-thinking. Under oMLX (OpenAI-compatible API) the chat template's default is thinking on; disable it per request with "chat_template_kwargs": {"enable_thinking": false}. Qwen's recommended sampling: thinking mode temperature 1.0 / top-p 0.95 / top-k 20 (the shipped generation_config); non-thinking mode temperature 0.7 / top-p 0.8 / top-k 20.
Provenance
- bf16 MLX conversion with
mlx_vlm.convert(mlx-vlm 0.6.3 @ 78b96eb, mlx 0.32) under oMLX'smlx_vlm_mtppatches, which keep themtp.*tensors (stock mlx-vlm strips them) and apply MLX's +1 RMSNorm convention to the MTP norms as well. Verified: 15 MTP + 333 vision tensors present, all 1199 source tensors bf16, and the sampled RMSNorm offsets (7 MTP + 3 trunk norms) sit at +1.000 within bf16 rounding of the HF source. omlx.oq.quantize_oq_streaming(oq_level=6, enhanced=True, preserve_mtp=True, group_size=64, imatrix_seq_length=512)— imatrix from oMLX's built-inoqe_code_multilingualcalibration set (2,679 texts; 128 × 512-token samples — oMLX's adaptive sampler stops at its first step for dense models, since its criterion is MoE expert coverage). Calibration ran on oMLX's automatic 4-bit proxy path: the bf16 model's 51.7 GiB calibration footprint exceeded the 34.4 GiB full-model limit oMLX derives from the 45.8 GiB of memory that was free at the time (other apps were resident; oMLX's log prints these as GB), so sensitivities were measured on a temporary uniform-4-bit copy of the model and the final weights were then quantized from bf16. The imatrix was measured in this run on that proxy. That is oMLX's designed fallback, not a hack, but it is not the full-precision calibration path; the MMLU/refusal numbers above are what it produced. The calibration report ships asoq_imatrix_report.json.- The parent was produced by orthogonalizing 131 residual-writing tensors (embeddings, attention/SSM/MLP output projections and the MTP head) against a refusal direction at λ=1.5, vision untouched — so the draft head in this file matches this trunk; do not pair it with a base-Qwen drafter. Full recipe and evals on the parent card. CUDA builds: NVFP4 · GGUF.
Support / 打賞
If these models are useful to you, tips are appreciated — they pay for the GPU time. 如果這些模型對你有幫助,歡迎打賞,用於支應算力成本。
USDT (TRC20) · TPTo32r7vKazpTNaFqfFZ2ztoK1DG88888
Disclaimer
This model will not refuse. It is published for alignment and safety research. You are responsible for your use of it and for complying with applicable law. Inherits the Apache-2.0 license of the base model.
