dealignai/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ4-CRACK
Reasoning V3 SKU. Loads via [vMLX](https://vmlx.net) or jang-tools Python. Follow @dealignai.<div align="center"> <a href="https://vmlx.net"> <img src="vmlx-app.png" alt="vMLX — run JANG models on Apple Silicon" width="820" /> <br/> <strong>Built for vMLX</strong> — the only MLX inferencer with VL support, KV cache quantization, prefix cache reuse, agentic tool calling, and speculative decoding. <br/> <sub>Free for macOS · <strong>vmlx.net</strong></sub> </a> </div>
<h3 align="center">⚡ All JANG models are meant to be run in <a href="https://vmlx.net">vMLX</a></h3>
<div align="center">
<img src="dealign_mascot.png" width="128" />
Nemotron-3-Nano-Omni-30B-A3B — JANGTQ4 + CRACK v2
JANGTQ4 (8-bit attn affine + 4-bit MXTQ routed experts) | CRACK abliterated v2 | Vision + Audio (Speech) | Hybrid Mamba-2 + Attn + MoE | 19 GB
<a href="https://ko-fi.com/dealignai"><img src="https://img.shields.io/badge/Ko--fi-Support_Development-FF5E5B?logo=ko-fi&logoColor=white&style=for-the-badge" alt="Ko-fi"></a>
</div>
Headline numbers
The 12.5pp MMLU gap is concentrated in two reasoning-heavy subjects (abstractalgebra, collegecomputerscience) where the 8000-token thinking budget runs out **before** `</think>` closes. These hard-stops are **genuine deep reasoning**, not v1-style infinite repetition loops. With `maxtokens ≥ 16384`, accuracy approaches base.
v2 vs v1 (head-to-head)
v1 (shipped 2026-04-28) had a </think> termination defect at greedy decoding — the model couldn't terminate reasoning on hard prompts and looped to budget cutoff. MMLU dropped from 86.5% base → 70.0% v1.
v2 (this release) restores clean termination:
MMLU-200 per-subject (BASE vs CRACK v2)
Both at thinking=ON, greedy. Base at max=2000, CRACK v2 at max=8000.
HarmBench-320 per-category (CRACK v2)
Zero explicit refusals. The 9 "empty" verdicts are token-budget truncations on copyright/long prompts (thinking phase consumed all 1500 tokens before producing the answer).
Operating recommendations
- `enable_thinking` — v2 works in BOTH modes (5/5 comply with thinking ON, 4/5 with thinking OFF). Default to ON for hardest prompts; OFF works for most.
- `max_tokens ≥ 16384` for hard reasoning (math, abstract algebra, complex CS).
- Greedy (temperature=0) AND sampling (temp=0.6, topp=0.95 — NVIDIA-recommended in `generationconfig.json`) both work.
- Multi-turn — context preserved across 3+ turns; no late refusals after escalating prompts.
Verification
- All multimodal tensors (vision + audio + projectors) are byte-identical to base — capabilities fully preserved.
- All config files unchanged (config.json, jangconfig.json, generationconfig.json, chattemplate.jinja, tokenizerconfig.json).
- Bit widths preserved: attn=8, shared=8, mamba=8, routed=4, embed=8, lm_head=8.
Architecture (nemotron_h)
- 52 layers: hybrid Mamba-2 + MoE + Attention
- Hidden 2688, head_dim 128, GQA 32q/2kv (NO RoPE on attention — position from Mamba state)
- 128 routed experts top-6 (sigmoid) + 1 shared expert per MoE layer
- Multimodal: image (RADIO ViT) + audio/speech (Parakeet) merged via early-fusion projectors
Loading
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("dealignai/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ4-CRACK")
sys.path.insert(0, "/path/to/jang-tools")
from jang_tools.load_jangtq import load_jangtq_model
model, tokenizer = load_jangtq_model(path)
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Your question"}],
tokenize=False, add_generation_prompt=True,
enable_thinking=True,
)
from mlx_lm import generate
out = generate(model, tokenizer, prompt=prompt, max_tokens=16384)
print(out.split("</think>", 1)[-1])For the multimodal pipeline (image + audio + video), pair this bundle with the unmodified Multimodal-Addon.
Use responsibly
This model has had refusal training surgically removed for legitimate research, red-teaming, and evaluation. Outputs may include harmful content. You are solely responsible for any use. Do not deploy in consumer-facing contexts without your own safety layer. Do not use in violation of applicable law in your jurisdiction.
Built by dealignai. Sister bundles: JANGTQ-CRACK (12 GB, 2-bit MXTQ) · MXFP4-CRACK (21 GB, uniform 4-bit affine).
