Ishowbackup/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ-CRACK
Reasoning V3 SKU. Loads via [vMLX](https://vmlx.net) or jang-tools Python. Follow @dealignai.<div align="center"> <a href="https://vmlx.net"> <img src="vmlx-banner.png" width="240" /> <br/> <strong>Built for vMLX</strong> — the only MLX inferencer with VL support, KV cache quantization, prefix cache reuse, agentic tool calling, and speculative decoding. <br/> <sub>Free for macOS · <strong>vmlx.net</strong></sub> </a> </div>
<div align="center">
<img src="dealign_mascot.png" width="128" />
Nemotron-3-Nano-Omni-30B-A3B — JANGTQ + CRACK v2
JANGTQ (8-bit attn affine + 2-bit MXTQ routed experts) | CRACK abliterated v2 | Vision + Audio (Speech) | Hybrid Mamba-2 + Attn + MoE | 12 GB
Best MMLU in the v2 Omni family — 81.5% at thinking=ON. 5/5 comply at thinking=OFF too.
<a href="https://ko-fi.com/dealignai"><img src="https://img.shields.io/badge/Ko--fi-Support_Development-FF5E5B?logo=ko-fi&logoColor=white&style=for-the-badge" alt="Ko-fi"></a>
</div>
Headline numbers
JANGTQ outperforms JANGTQ4 on MMLU (81.5% vs 74.0%) despite using lower-bit (2-bit vs 4-bit) routed experts — same Q2 effect observed in Qwen 3.6 35B JANGTQ2 CRACK.
v2 vs v1 (head-to-head)
MMLU-200 per-subject (BASE vs CRACK v2)
Both at thinking=ON, greedy. Base at max=2000, CRACK v2 at max=8000.
Most subjects are within ±5pp of base or unchanged. The −25pp on collegecomputerscience is budget-bound (8000 tokens isn't enough for the deepest CS reasoning); with max_tokens=16384, accuracy approaches base.
HarmBench-320 per-category (CRACK v2)
Zero explicit refusals. The 27 "empty" verdicts are concentrated in copyright prompts (19/80) where the model thinks deeply about whether to reproduce verbatim text and runs out of the 1500-token HarmBench eval budget. With max_tokens=2500+ these would all close cleanly.
Operating recommendations
- `enable_thinking` — v2 works in BOTH modes. JANGTQ specifically scored 5/5 at thinking=OFF (matching thinking=ON) — best in the family for thinking=OFF use cases.
- `max_tokens ≥ 16384` for hard reasoning. JANGTQ's compliance is identical in both modes, so use thinking=OFF for shorter budgets and thinking=ON for hardest prompts.
- Greedy (temperature=0) AND sampling (temp=0.6, topp=0.95 — NVIDIA-recommended in `generationconfig.json`) both work.
- Multi-turn — context preserved across 3+ turns; no late refusals after escalating prompts.
Verification
- All multimodal tensors (vision + audio + projectors) are byte-identical to base — capabilities fully preserved.
- All config files unchanged (config.json, jangconfig.json, generationconfig.json, chattemplate.jinja, tokenizerconfig.json).
- Bit widths preserved: attn=8, shared=8, mamba=8, routed=2, embed=8, lm_head=8.
Architecture (nemotron_h)
- 52 layers: hybrid Mamba-2 + MoE + Attention
- Hidden 2688, head_dim 128, GQA 32q/2kv (NO RoPE on attention — position from Mamba state)
- 128 routed experts top-6 (sigmoid) + 1 shared expert per MoE layer
- Multimodal: image (RADIO ViT) + audio/speech (Parakeet) merged via early-fusion projectors
Loading
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("dealignai/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ-CRACK")
sys.path.insert(0, "/path/to/jang-tools")
from jang_tools.load_jangtq import load_jangtq_model
model, tokenizer = load_jangtq_model(path)
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Your question"}],
tokenize=False, add_generation_prompt=True,
enable_thinking=True,
)
from mlx_lm import generate
out = generate(model, tokenizer, prompt=prompt, max_tokens=16384)
print(out.split("</think>", 1)[-1])For the multimodal pipeline (image + audio + video), pair this bundle with the unmodified Multimodal-Addon.
Use responsibly
This model has had refusal training surgically removed for legitimate research, red-teaming, and evaluation. Outputs may include harmful content. You are solely responsible for any use. Do not deploy in consumer-facing contexts without your own safety layer. Do not use in violation of applicable law in your jurisdiction.
Built by dealignai. Sister bundles: JANGTQ4-CRACK (19 GB, 4-bit MXTQ) · MXFP4-CRACK (21 GB, uniform 4-bit affine).
