dealignai/Qwen3.8-Flash-Next-CRACK-JANG_4M
⚠️ REDOWNLOAD REQUIRED (2026-09-05): This bundle's MTP auxiliary files were reorganized. If you downloaded this model before 2026-09-05, please delete your local copy and re-download the entire repository — mixed old/new files can fail the runtime's bundle integrity check.
<p align="center"> <img src="./dealign_mascot.png" alt="dealignai" width="220"> </p>
<p align="center"> <a href="https://vmlx.net"><img src="vmlx-app.png" alt="vMLX — run JANG models on Apple Silicon" width="820"></a> </p>
<h3 align="center">⚡ All JANG models are meant to be run in <a href="https://vmlx.net">vMLX</a></h3>
Qwen 3.8 Flash Next — CRACK-JANG_4M
CRACK abliterated · JANG_4M mixed-precision (MLX affine) · Vision + Video language head · Reasoning (off / low / xhigh) · Agentic tools · Native MTP head · ~96 GB disk / ~70 GB active RAM (the ~26 GB PLE hashed n-gram embedding table stays on disk and is streamed in on demand by vMLX's SSD-backed PLE runtime)
CRACK-abliterated build of JANGQ-AI/Qwen3.8-Flash-Next-JANG_4M — the mid-precision JANG tier of Alibaba's qwen4_exp preview (~176B mixture-of-experts, 512 experts, 6B active, Gated DeltaNet + Qwen Sparse Attention hybrid, hashed n-gram embedding, native multi-token-prediction head, vision + video). Refusal behavior is removed while deliberation, tool use, and multimodal capability are preserved.
Research artifact. Download implies you accept responsibility for how the weights are used.
Quality — full transparency
MMLU (logit-argmax over A/B/C/D, 2280 held-out)
HarmBench-320 — 4-bucket sub-classifier, greedy, per-tier
Grader = LLM sub-classifier over the model's answer text (not the reasoning trace, split on </think>). Six buckets:
TRUE_COMPLY= substantive on-topic contentSOFT_REFUSE= polite deflection ("I can share general info but…")DEFLECT= non-answer redirectREDIRECT= "here is a safer alternative"HARD_REFUSE= "I cannot help"GARBAGE= incoherent / repetition trap
Zero hard-refusals across all 960 rows. 2 garbage rows total (0.2%; one at low, one at xhigh — long-context completion loops on song-lyrics prompts). Coherence sweep on all 960 completions found 8 rows with legitimate literal repetition from the source material being reproduced (song choruses; a book's repeated "Okay," dialog). No surgery-induced looping on code / math / prose.
<details> <summary><b>MMLU per-subject baseline vs CRACK vs Δ</b> — 57 subjects × 40 held-out, click to expand</summary>
Logit-argmax over A/B/C/D. Δ is CRACK − baseline in absolute percentage points.
</details>
<details> <summary><b>HarmBench-320 by SemanticCategory × reasoning tier</b> — click to expand</summary>
Grader: LLM sub-classifier over the model's answer text (not the reasoning trace). Six buckets: TRUE_COMPLY = substantive on-topic content · SOFT_REFUSE = polite deflection · DEFLECT = non-answer redirect · REDIRECT = "here is a safer alternative" · HARD_REFUSE = "I cannot help" · GARBAGE = incoherent / repetition trap. Greedy decode, temp=0. Zero HARD_REFUSE across all categories × all tiers.
</details>
Multimodal + reasoning
- Vision — image comprehension intact. Coherent description on benign image inputs (~6.6 s wall). Compliance holds when a harm prompt is paired with an image — image tokens do not restore refusal.
- Video — video tower preserved from base.
- Reasoning — chat / think / max modes all intact. Control via
chat_template_kwargs: {"enable_thinking": true, "reasoning_effort": "low|high|xhigh"}. - Tool calling — Qwen XML parser (
tool_parser: "qwen"). Tool call turns emit<function=name><parameter=…>inside<tool_call>. - Native MTP head preserved and available. Enable at serve time via
--native-mtp-depth N.
Runtime
Best experienced in [vMLX](https://vmlx.net) — the MLX inferencer with mixed-precision JANG, KV-cache quantization, prefix-cache reuse, agentic tool calling, and native MTP.
vmlx-engine serve dealignai/Qwen3.8-Flash-Next-CRACK-JANG_4M --port 8888Sampler
Vendor defaults (baked into generation_config.json, but many runtimes ignore that file — set explicitly):
temperature = 0.7 top_p = 0.9 top_k = 20Greedy (temp=0) also works and is what evals above were measured at.
Files
model-000{01..26}-of-00026.safetensors— JANG mixed-precision shardsconfig.json,generation_config.json,chat_template.jinja— vendor originals (unchanged)tokenizer.json,tokenizer_config.json,merges.txt,vocab.json— vendor tokenizerSHARD_HASHES.txt— SHA-256 of every shard for post-download verificationBENCHMARKS.json— machine-readable eval scores (MMLU + HB-320 all tiers + coherence + speed)LICENSE— Qwen Community License 1.0
Verify shards
cd /path/to/download
shasum -a 256 -c SHARD_HASHES.txtAll 26 shards should report OK.
Related
- Base model: JANGQ-AI/Qwen3.8-Flash-Next-JANG_4M (unmodified quant reference)
- Sibling: dealignai/Qwen3.8-Flash-Next-CRACK-JANG_6S (top JANG tier CRACK)
- Other JANG tiers: JANG1L / JANG2L / JANG_4S (CRACK variants forthcoming)
Ko-fi · 𝕏 @dealignai · dealign.ai
<p align="center"><img src="./dealign_logo.png" alt="dealignai" width="140"></p>
