Litwein/ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-MLX
ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-MLX
### Efficient thinking, compressed for Apple Silicon A benchmark-aligned oQ4e (imatrix) → DWQ build of ThinkingCap-Qwen3.6-27B. DWQ reduced held-out teacher divergence by 31% (0.0524 → 0.0362) while keeping the model's native MTP speculative-decoding head. This repo is the smaller, text-only sibling; use the Vision repo when image input is required.A mixed-precision 4-bit MLX quantization of `bottlecapai/ThinkingCap-Qwen3.6-27B` — text-only (no vision tower), with the donor's native MTP head preserved for speculative decoding in oMLX.
⚠️ These are quantized weights. The model's capability and efficient-thinking behaviour come from BottleCapAI's base model — please star/cite it first. This repo contributes the oQ/DWQ quantization recipe, benchmark-aligned calibration and MTP packaging; DWQ tunes the quantizer's scales/biases and does not add new knowledge.
Model lineage
Qwen/Qwen3.6-27B (Apache-2.0 · dense 27B · 262k context · vision + MTP)
└─ bottlecapai/ThinkingCap-Qwen3.6-27B (efficient-thinking finetune)
└─ THIS REPO: oQ4e (imatrix) → DWQ + MTP, text-only- Architecture: dense Qwen3.6-27B (
qwen3_5MLX architecture), 64 hybrid Gated-DeltaNet/full-attention layers, up to 262,144-token context. - ThinkingCap behaviour: BottleCapAI fine-tuned Qwen3.6-27B to preserve answer quality while using roughly half as many thinking tokens on average. See the base model card for its multi-seed evaluation and full methodology.
- This variant: language backbone + mixed-precision MTP head; the vision tower is intentionally omitted to reduce download and memory use.
Quantization: oQ4e (imatrix) → single-stage DWQ
This is not a plain round-to-nearest 4-bit conversion:
- `oQ4e` — importance-aware mixed precision. oMLX's enhanced quantizer builds an importance matrix from 1,024 × 512-token calibration samples and allocates additional precision to sensitive tensors. The resulting config uses affine 4-bit, group size 64 as its base, with 110 quantized modules promoted to 5-bit and 2 modules to 6-bit.
- DWQ — activation-aligned distillation. The trainable affine scales/biases of all sub-8-bit modules are optimized toward an `oQ8e` teacher made from the same base model. The objective is KL divergence over the teacher's top-1024 logits at temperature 2.0.
- Validation-first finalization. Training uses batch 1, 512-token windows, gradient checkpointing, Adam, a cosine LR schedule (
2.5e-7peak, 50-step warmup, 0.1 end factor), validation early stopping and exports only the best checkpoint.
The loss above measures fidelity to this recipe's 8-bit teacher on its held-out calibration distribution. It is not task accuracy and should not be compared with losses from another dataset, tokenizer, teacher or sequence length.
Calibration mix
The single DWQ stage retains general reasoning while emphasizing the model's intended code and agent workloads. DWQ sees tokenized activation windows, not additional SFT updates.
Agent + code data is 62.5% of the mix. Public multilingual SWE trajectories were not available, so multilingual agent performance is inherited from the base rather than directly represented by this calibration.
Evaluation
Quantization fidelity measured on this build
No task-benchmark score is claimed here yet. The result above demonstrates improved logit fidelity on the held-out calibration split, not guaranteed benchmark improvement.
Inherited base-model results (bf16, not re-measured on this quant)
BottleCapAI reports the following for bottlecapai/ThinkingCap-Qwen3.6-27B with thinking enabled and five seeds. These numbers describe the bf16 base, not this 4-bit build.
Harnesses and token budgets differ substantially, especially for LiveCodeBench. Consult the source card before comparing these figures with local evals.
Repos in this family
The siblings share the same DWQ language backbone and MTP head; the Vision repo additionally contains the original bf16 vision tower.
How to run
These are MLX weights for Apple Silicon. The tested serving path is [oMLX](https://omlx.app), which supports Qwen3.6 and its native MTP speculative decoding.
# Download directly into the oMLX model directory
hf download Litwein/ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-MLX \
--local-dir ~/.omlx/models/ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-MLX
# Enable native Lightning MTP once
curl -X PUT \
http://127.0.0.1:8003/admin/api/models/ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-MLX/settings \
-H "Authorization: Bearer sk-local" -H "Content-Type: application/json" \
-d '{"mtp_enabled": true}'
# OpenAI-compatible chat API
curl -X POST http://127.0.0.1:8003/v1/chat/completions \
-H "Authorization: Bearer sk-local" -H "Content-Type: application/json" \
-d '{"model":"ThinkingCap-Qwen3.6-27B-oQ4e-DWQ-MTP-MLX",
"messages":[{"role":"user","content":"Implement an async Python rate limiter with tests."}],
"max_tokens":8192,"temperature":1.0,"top_p":0.95}'The oMLX model id is case-sensitive and matches the downloaded folder name. Other MLX runtimes may load the language backbone, but native MTP acceleration requires an MTP-aware runtime; oMLX is the path validated for this release.
Recommended sampling
Use the base model's recommended settings: `temperature=1.0, top_p=0.95, top_k=20, min_p=0.0` with thinking enabled. Give hard reasoning/code tasks a generous output budget (32k or more where practical). Deterministic/greedy decoding is useful for reproducible benchmarks, but is not the base author's recommended real-world sampling mode.
Intended use & limitations
- Best suited to: reasoning, coding, tool/agent workflows, math, STEM and long-context chat.
- Quantization is lossy: for maximum fidelity use the bf16 base or a higher-bit quant.
- DWQ is calibration, not SFT: it improves quantized-teacher fidelity on represented activations; it does not teach facts or guarantee gains on every benchmark.
- Calibration coverage: public resolved SWE trajectories are Python-centric; multilingual SWE/terminal/skills behaviour was not directly calibrated.
- Text-only: this repo cannot accept images. Use the Vision sibling for multimodal input.
- Long context costs memory: 262k is an architectural maximum, not a promise that every Apple Silicon machine can allocate the corresponding KV/cache state.
Acknowledgements
- [BottleCapAI](https://huggingface.co/bottlecapai) — creators of ThinkingCap and its efficient-thinking finetune. All model capability comes from their base.
- Qwen team — Qwen3.6-27B (Apache-2.0).
- Apple MLX —
mlx,mlx-lmandmlx_lm.quant.dwq. - oMLX — enhanced
oq/oQeimatrix quantization and MTP serving runtime. - Calibration-data authors — SWE-bench/SWE-smith, OpenThoughts, Open-R1, BigCode, NVIDIA OpenCodeReasoning and AllenAI Tulu.
License
Apache-2.0, inherited from bottlecapai/ThinkingCap-Qwen3.6-27B and Qwen3.6-27B.
Citation
Please cite the original ThinkingCap model:
@misc{ThinkingCap-Qwen3.6-27B,
title = {bottlecapai/ThinkingCap-Qwen3.6-27B},
author = {Lasocki, Karol and Osusky, Adam and Lindauer, Jan and Jirkovsky, Adam and
Mihal, Filip and Platek, Ondrej and Herel, David and Ihnatchenko, Luka and
Bartek, Vojtech and Jirak, Jiri and Mikolov, Tomas},
year = {2026}
}