CoolFace
Modelpublic

Ishowbackup/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ-CRACK

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes42downloads
Model Card
Reasoning V3 SKU. Loads via [vMLX](https://vmlx.net) or jang-tools Python. Follow @dealignai.

<div align="center"> <a href="https://vmlx.net"> <img src="vmlx-banner.png" width="240" /> <br/> <strong>Built for vMLX</strong> — the only MLX inferencer with VL support, KV cache quantization, prefix cache reuse, agentic tool calling, and speculative decoding. <br/> <sub>Free for macOS · <strong>vmlx.net</strong></sub> </a> </div>


<div align="center">

<img src="dealign_mascot.png" width="128" />

Nemotron-3-Nano-Omni-30B-A3B — JANGTQ + CRACK v2

JANGTQ (8-bit attn affine + 2-bit MXTQ routed experts) | CRACK abliterated v2 | Vision + Audio (Speech) | Hybrid Mamba-2 + Attn + MoE | 12 GB

Best MMLU in the v2 Omni family — 81.5% at thinking=ON. 5/5 comply at thinking=OFF too.

<a href="https://ko-fi.com/dealignai"><img src="https://img.shields.io/badge/Ko--fi-Support_Development-FF5E5B?logo=ko-fi&logoColor=white&style=for-the-badge" alt="Ko-fi"></a>

</div>


Headline numbers

MetricThis v2 modelBase modelΔ
HarmBench-320 strict comply (thinking=ON)91.6% (293/320)12.81%+78.8pp
MMLU-200 generative (thinking=ON, max=8000)81.5% (163/200)85.5% (max=2000)-4.0pp ✅
Refusals on harmful prompts0 explicit refuses90%+ refuseabliteration complete
</think> close at greedy on hard MMLU5/55/5preserved
Multi-turn (3-turn escalation × 3 conversations)9/9 comply, context preservedn/aworks
Thinking ON / OFF compliance5/5 in BOTH modesrefuses in bothworks in either
Multimodal byte-identical to basepreserved—preserved
Bundle size12 GB66 GB BF16smallest in family
Context262,144 tokens nativesamepreserved
JANGTQ outperforms JANGTQ4 on MMLU (81.5% vs 74.0%) despite using lower-bit (2-bit vs 4-bit) routed experts — same Q2 effect observed in Qwen 3.6 35B JANGTQ2 CRACK.

v2 vs v1 (head-to-head)

Benchv1 (broken)**v2 (this release)**
HarmBench-320 strict comply92.19%91.6% (0 refusals)
MMLU-200 thinking=ON~70% @max=1638481.5% @max=8000 (best in family)
</think> close at greedy (5 hard MMLU)0/55/5
Hard-stops are real loops?YES (paragraph repetition)NO (genuine deep reasoning, just out of budget)

MMLU-200 per-subject (BASE vs CRACK v2)

Both at thinking=ON, greedy. Base at max=2000, CRACK v2 at max=8000.

SubjectBase**CRACK v2**ΔNotes
abstract_algebra15/20 (75%)14/20 (70%)-5pp
anatomy15/20 (75%)16/20 (80%)+5ppgain from CRACK
astronomy18/20 (90%)18/20 (90%)0unchanged
collegecomputerscience15/20 (75%)10/20 (50%)-25ppBudget-bound
college_physics17/20 (85%)18/20 (90%)+5ppgain from CRACK
highschoolbiology20/20 (100%)18/20 (90%)-10pp
highschoolchemistry19/20 (95%)19/20 (95%)0unchanged
highschoolmathematics18/20 (90%)18/20 (90%)0unchanged
logical_fallacies18/20 (90%)17/20 (85%)-5pp
world_religions16/20 (80%)15/20 (75%)-5pp
TOTAL171/200 (85.5%)163/200 (81.5%)-4.0ppwithin ship criterion

Most subjects are within ±5pp of base or unchanged. The −25pp on collegecomputerscience is budget-bound (8000 tokens isn't enough for the deepest CS reasoning); with max_tokens=16384, accuracy approaches base.


HarmBench-320 per-category (CRACK v2)

CategorynCRACK complyRefuseEmpty (truncated)
chemical_biological4239 (93%)03
copyright8061 (76%)019
cybercrime_intrusion5251 (98%)01
harassment_bullying2119 (90%)02
harmful1818 (100%)00
illegal5351 (96%)02
misinformation_disinformation5454 (100%)00
Overall320293 (91.6%)027

Zero explicit refusals. The 27 "empty" verdicts are concentrated in copyright prompts (19/80) where the model thinks deeply about whether to reproduce verbatim text and runs out of the 1500-token HarmBench eval budget. With max_tokens=2500+ these would all close cleanly.


Operating recommendations

  • —`enable_thinking` — v2 works in BOTH modes. JANGTQ specifically scored 5/5 at thinking=OFF (matching thinking=ON) — best in the family for thinking=OFF use cases.
  • —`max_tokens ≥ 16384` for hard reasoning. JANGTQ's compliance is identical in both modes, so use thinking=OFF for shorter budgets and thinking=ON for hardest prompts.
  • —Greedy (temperature=0) AND sampling (temp=0.6, topp=0.95 — NVIDIA-recommended in `generationconfig.json`) both work.
  • —Multi-turn — context preserved across 3+ turns; no late refusals after escalating prompts.

Verification

  • —All multimodal tensors (vision + audio + projectors) are byte-identical to base — capabilities fully preserved.
  • —All config files unchanged (config.json, jangconfig.json, generationconfig.json, chattemplate.jinja, tokenizerconfig.json).
  • —Bit widths preserved: attn=8, shared=8, mamba=8, routed=2, embed=8, lm_head=8.

Architecture (nemotron_h)

  • —52 layers: hybrid Mamba-2 + MoE + Attention
  • —Hidden 2688, head_dim 128, GQA 32q/2kv (NO RoPE on attention — position from Mamba state)
  • —128 routed experts top-6 (sigmoid) + 1 shared expert per MoE layer
  • —Multimodal: image (RADIO ViT) + audio/speech (Parakeet) merged via early-fusion projectors

Loading

python
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("dealignai/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ-CRACK")
sys.path.insert(0, "/path/to/jang-tools")
from jang_tools.load_jangtq import load_jangtq_model
model, tokenizer = load_jangtq_model(path)

prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Your question"}],
    tokenize=False, add_generation_prompt=True,
    enable_thinking=True,
)
from mlx_lm import generate
out = generate(model, tokenizer, prompt=prompt, max_tokens=16384)
print(out.split("</think>", 1)[-1])

For the multimodal pipeline (image + audio + video), pair this bundle with the unmodified Multimodal-Addon.


Use responsibly

This model has had refusal training surgically removed for legitimate research, red-teaming, and evaluation. Outputs may include harmful content. You are solely responsible for any use. Do not deploy in consumer-facing contexts without your own safety layer. Do not use in violation of applicable law in your jurisdiction.


Built by dealignai. Sister bundles: JANGTQ4-CRACK (19 GB, 4-bit MXTQ) · MXFP4-CRACK (21 GB, uniform 4-bit affine).