CoolFace
Modelpublic

airagrp/ThinkingCap-Qwen3.8-27B-mlx-nvfp4

sourceHugging Faceotherupdated 2d agoView on Hugging Face
0likes27downloads
Model Card

This repository contains `bottlecapai/ThinkingCap-Qwen3.8-27B` converted to MLX format with a mixed-precision quantization recipe, using mlx-vlm 0.6.17.

Quantization recipe

ModulePrecision
MLP gate_proj / up_proj / down_proj (64 layers)nvfp4 (group_size=16, bits=4)
Full attention q_proj / k_proj / v_proj / o_proj (16 layers)mxfp8 (group_size=32, bits=8)
Linear (GDN) attention in_proj_* / out_proj (48 layers)mxfp8 (group_size=32, bits=8)
Token embeddings (embed_tokens)bfloat16
Output head (lm_head)bfloat16
MTP headbfloat16
Vision towerbfloat16
  • —Effective size: ~24 GB (7.9 bits per weight), base model is ~54 GB in bfloat16.
  • —Quantized modules: mxfp8 (groupsize=32, bits=8) / nvfp4 (groupsize=16, bits=4); bfloat16 modules are stored as-is. Per-module precision is detected from the presence of .scales tensors.

MTP

The native MTP head is merged into this checkpoint as language_model.mtp.* tensors (15 tensors, bfloat16, norms in the MLX +1 convention), stored in mtp.safetensors and referenced from model.safetensors.index.json — it is not a separate drafter model. Use it for speculative decoding (--draft-kind mtp in mlx-vlm) or ignore it; base inference is unaffected.

Use with mlx-vlm

bash
pip install mlx-vlm
python
import mlx_vlm

model, processor = mlx_vlm.load("airagrp/ThinkingCap-Qwen3.8-27B-mlx-nvfp4")
response, _ = mlx_vlm.generate(
    model,
    processor,
    prompts="In one sentence, what is MLX?",
    max_tokens=64,
)
print(response)
bash
mlx_vlm.generate --model airagrp/ThinkingCap-Qwen3.8-27B-mlx-nvfp4 --prompt "In one sentence, what is MLX?" --max-tokens 64

Use with MLX directly

Load with the standard MLX safetensors layout; quantized weights use mxfp8 (groupsize=32, bits=8) / nvfp4 (groupsize=16, bits=4).

Citations / license

This checkpoint derives from `bottlecapai/ThinkingCap-Qwen3.8-27B` (a finetune); the finetune weights are licensed under polyform-small-business-1.0.0. See LICENSE in this repo. Refer to the original model card for architecture details, benchmarks, and usage guidelines.

Quality benchmarks

Perplexity and KL divergence on the WikiText-2 raw test split (297,193 tokens, 2048-token windows, fresh KV cache per window), measured against the bf16 reference (ThinkingCap-Qwen3.8-27B).

ModelSize (GiB)PPL (lower = better)ΔPPL vs bf16KLD vs bf16 (nats/token, lower = better)
ThinkingCap-Qwen3.8-27B (bf16 reference)51.757.008——
`ThinkingCap-Qwen3.8-27B-mlx-nvfp4`22.317.075+0.0680.0427
`ThinkingCap-Qwen3.8-27B-mlx-mxfp8`29.786.962-0.0460.0077
  • —KLD = mean per-token D_KL(p_bf16 ‖ p_model) over the full vocabulary — how much the quantized model's next-token distribution drifts from bf16.
  • —Data: wikitext-2-raw-v1 test split, SHA-256 696cca6b65a171b0…; tokenizer: ThinkingCap-Qwen3.8-27B; mlx-vlm 0.6.17.
  • —Benchmarked 2026-09-24 with benchmark_ppl_kld.py (2048-token streaming windows, all positions except the first of each window scored).

Plots

[image]

[image]