JANGQ-AI/Nemotron-3-Super-120B-A12B-JANG_4M
<p align="center"> <a href="https://mlx.studio"><img src="https://raw.githubusercontent.com/jjang-ai/jangq/main/assets/mlx-studio-light.png" alt="MLX Studio" width="500"></a> </p>
<p align="center"> <a href="https://mlx.studio"><img src="https://mlx.studio/assets/screenshots/mlx-studio-featured.png?v=1" alt="MLX Studio App" width="600"></a> </p>
<h4 align="center"><a href="https://mlx.studio">MLX Studio</a> — the only app that natively supports JANG models with reasoning</h4>
93% MMLU at same size as MLX 4-bit. JANG_4M matches MLX 4-bit quality with 8-bit attention protection. Hybrid Mamba-2 SSM + Latent MoE + Attention.
LM Studio, Ollama, oMLX do NOT support JANG format. Use [MLX Studio](https://mlx.studio) or \pip install "jang[mlx]>=2.1.5"\.<p align="center"> <img src="https://raw.githubusercontent.com/jjang-ai/jangq/main/assets/jangq-logo-dark.png" alt="JANG" width="300"> </p>
<p align="center"> <a href="https://vmlx.net"><img src="vmlx-app.png" alt="vMLX — run JANG models on Apple Silicon" width="820"></a> </p>
<h3 align="center">⚡ All JANG models are meant to be run in <a href="https://vmlx.net">vMLX</a></h3>
<h3 align="center">Nemotron-3-Super-120B-A12B — JANG_4M (4.1-bit, 8-bit attention) — Reasoning</h3> <p align="center"><b>JANG</b> — Jang Adaptive N-bit Grading | The GGUF Equivalent for MLX</p>
<p align="center"> <a href="https://github.com/jjang-ai/jangq"><img src="https://img.shields.io/badge/GitHub-Source_Code-blue?logo=github" alt="GitHub"></a> <a href="https://pypi.org/project/jang/"><img src="https://img.shields.io/pypi/v/jang?label=PyPI&color=green" alt="PyPI"></a> <a href="https://jangq.ai"><img src="https://img.shields.io/badge/Web-jangq.ai-orange" alt="Website"></a> <a href="https://x.com/dealignai"><img src="https://img.shields.io/badge/X-@dealignai-black?logo=x" alt="X/Twitter"></a> </p>
JANG is fully open-source. Quantization engine, research, and full commit history: github.com/jjang-ai/jangq. Created by Jinho Jang.
Key Features
- 93.0% MMLU (200 questions, reasoning mode) — matches MLX 4-bit at same size
- 55.1 tok/s generation, 154 tok/s prefill
- 63 GB on disk, 61.2 GB GPU RAM
- Reasoning mode: \
<think>...</think>\step-by-step problem solving - Hybrid architecture: 40 Mamba-2 SSM + 40 Latent MoE (512 experts) + 8 Dense Attention layers
- bfloat16 compute: auto-detected for 512-expert models
Results: JANG vs MLX (200-question MMLU)
Per-subject comparison. All models tested with and without reasoning using identical methodology.
Summary
JANG4M nearly ties MLX 4-bit (93.0% vs 93.5%) at the same 63 GB size with 8-bit attention protection. MLX 3-bit cannot be created — \`mlxlm.convert\` crashes on Nemotron's mtp.* weights. Only JANG can produce sub-4-bit quantizations.
Also see: JANG_2L (43 GB) — 20 GB smaller, fits 64 GB Macs, 75% no-think / 86% reasoning.
Specs
Requirements
- Apple Silicon Mac with 64+ GB unified memory
- [MLX Studio](https://mlx.studio) or \
pip install "jang[mlx]>=2.1.5"\
Quick Start
\\\bash pip install "jang[mlx]>=2.1.5" \\\
\\\`python from jangtools.loader import loadjangmodel from mlxlm import generate
model, tokenizer = loadjangmodel("JANGQ-AI/Nemotron-3-Super-120B-A12B-JANG_4M")
With reasoning
messages = [{"role": "user", "content": "Explain quantum computing."}] prompt = tokenizer.applychattemplate(messages, tokenize=False, addgenerationprompt=True, enablethinking=True) result = generate(model, tokenizer, prompt=prompt, maxtokens=2048)
Without reasoning (faster)
prompt = tokenizer.applychattemplate(messages, tokenize=False, addgenerationprompt=True, enablethinking=False) result = generate(model, tokenizer, prompt=prompt, maxtokens=100) \\\`
Technical Notes
- Latent MoE: Nemotron-H compresses hidden states 4096→1024 before expert routing. JANG loader handles this automatically.
- bfloat16: Auto-detected for 512-expert models. Prevents float16 overflow. Zero quality impact.
- trust_remote_code: Custom Python files included (modelingnemotronh.py, configurationnemotronh.py).
<p align="center"> <b>JANG</b> — Created by <a href="https://jangq.ai">Jinho Jang</a> (eric@jangq.ai) · <a href="https://x.com/dealignai">@dealignai</a><br> <a href="https://github.com/jjang-ai/jangq">GitHub</a> · <a href="https://pypi.org/project/jang/">PyPI</a> · <a href="https://huggingface.co/JANGQ-AI">HuggingFace</a> </p>
한국어
Nemotron-3-Super-120B JANG_4M — MLX 4-bit과 동일한 크기(63 GB)에서 93% MMLU 달성.
\\\bash pip install "jang[mlx]>=2.1.5" \\\
