JANGQ-AI/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ4
<p align="center"> <a href="https://vmlx.net"><img src="vmlx-app.png" alt="vMLX — run JANG models on Apple Silicon" width="820"></a> </p>
<h3 align="center">⚡ All JANG models are meant to be run in <a href="https://vmlx.net">vMLX</a></h3>
Nemotron-3-Nano-Omni-30B-A3B-Reasoning · JANGTQ4
19.9 GB · ~82 tok/s decode on M4 Max · 30B / 3B-active hybrid Mamba-2 + Attention + MoE · native MLX, zero PyTorch in the hot path
Full multimodal (text + image + audio + video) port of NVIDIA's Nemotron-3-Nano-Omni-30B-A3B-Reasoning to Apple MLX, all four modalities running natively on Metal:
Bundle contents (single repo, everything included)
- LLM — 52-layer hybrid Mamba-2 + Attention + MoE at JANGTQ4 quantization (TurboQuant 4-bit codec on routed experts)
- Vision tower — NVIDIA RADIO ViT-Huge (32 blocks, 1280 hidden, 10 cls/register tokens) at fp16 — 1.31 GB
- Vision projector (mlp1) — LayerNorm + Linear + GELU + Linear → LLM hidden — 0.32 GB
- Sound encoder — parakeet (24-layer Conformer, full Transformer-XL relative-position attention) at fp16 — 1.22 GB
- Sound projector — RMSNorm + Linear + SquaredReLU + Linear → LLM hidden — 0.03 GB
- Source `.py` files — modeling.py, audio_model.py, image/video/audio processors (PyTorch fallback path)
- Codec sidecar (
jangtq_runtime.safetensors) — codebook + Hadamard signs (JANGTQ variants only)
Quantization recipe
Routed experts at 4-bit JANGTQ (Hadamard-rotated codebook, 16-entry centroid table per in_features × Lloyd-Max-fitted norms). Attention Q/K/V/O, shared experts, Mamba in/outproj, embed, lmhead at 8-bit affine. Mamba 1-D + router gate fp16. Codec sidecar jangtq_runtime.safetensors (codebook + Hadamard signs, seed=42) ships with the bundle. Loads via jang_tools.load_jangtq.load_jangtq_model().
All bundles in this family
Install
pip install jang_tools mlx mlx_lm pillow soundfile scipy librosa imageio[ffmpeg](Optional, for the PyTorch hybrid fallback only: pip install transformers torch torchaudio timm open_clip_torch)
Native MLX multimodal (recommended — zero PyTorch dependency)
import mlx.core as mx
from jang_tools.nemotron_omni.model import NemotronHOmni
chat = NemotronHOmni("JANGQ-AI/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ4", dtype=mx.float32)
# Text only
print(chat.turn("Capital of France?")) # "Paris."
# Image input — RADIO ViT runs natively in MLX on Metal (1.4s)
print(chat.turn("What's in this image?", images=["cat.jpg"]))
# Audio input — parakeet Conformer encoder native MLX (2.1s for 20s clip)
print(chat.turn("Transcribe what was said.", audio="speech.wav"))
# Video input — frame extraction + RADIO video_embedder + EVS pruning native MLX (3.6s for 8 frames)
print(chat.turn("Describe what happens.", video="clip.mp4",
video_target_frames=8, video_apply_evs=True))
# Mixed modality
print(chat.turn(
"Compare the image with the spoken description.",
images=["scene.jpg"], audio="description.wav",
))
# Multi-turn — KV + Mamba state persists across turns
print(chat.turn("And what about the previous image?")) # references prior turn
chat.reset() # new conversationReasoning ON / OFF
# Reasoning ON (default for Reasoning SKU): emits <think>...</think> + answer
chat.turn("Solve: 17 + 28 = ?", enable_thinking=True)
# Reasoning OFF: faster, more direct
chat.turn("Solve: 17 + 28 = ?", enable_thinking=False)Text-only fast path (mlxlm or loadjangtq)
For chat-only use cases, skip the multimodal load and use the LLM directly:
from mlx_lm import load, generate
model, tokenizer = load("JANGQ-AI/Nemotron-3-Nano-Omni-30B-A3B-JANGTQ4")
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content": "Capital of France?"}],
tokenize=False, add_generation_prompt=True,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=20)) # "Paris."(For JANGTQ4 / JANGTQ2: replace mlx_lm.load with from jang_tools.load_jangtq import load_jangtq_model. Vision/sound weights are silently dropped on the text-only path.)
Architecture (52 hybrid layers)
hybrid_override_pattern (52 chars):
"MEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEMEM*EMEMEMEME"
23 × M = Mamba-2 SSM (state-space, O(1) cache per token)
23 × E = MoE (128 routed × 6) (ReLU² activation, no gate_proj)
6 × * = Attention (GQA 32q / 2kv heads, NO RoPE, head_dim=128)
Multimodal towers (fp16, native MLX):
vision_model → RADIO ViT (NVIDIA C-RADIOv2-H)
mlp1 → LayerNorm + Linear + GELU + Linear → llm_hidden
sound_encoder → ParakeetEncoder (24 Conformer layers)
sound_projection → RMSNorm + Linear + SquaredReLU + Linear → llm_hidden
Cache (multi-turn):
M layers → MambaCache (size=2: conv state + ssm state) O(1)/token
* layers → KVCacheSimple O(L)/token
E layers → stateless
Native context: 262 144 tokens (no RoPE extrapolation needed)Special tokens
Sampling guidance
Swift / vMLX support
Native Swift port is in `vmlx-swift-lm` under Libraries/MLXVLM/Models/NemotronHOmni/. The full multimodal pipeline (NemotronHOmni wrapper + RADIOVision + Parakeet + Projectors + image/audio/video preprocessors) compiles cleanly. Shared video utilities in Libraries/MLXVLM/VLMVideoUtils.swift are reused by Qwen 2/2.5/3/3.5/3.6 VL and Kimi VL for cross-VLM compatibility.
import MLXVLM
let frames = try await vlmExtractFramesUniform(url: videoURL, targetFrames: 32)
let pixels = vlmStackFramesIntoChannels(frames, imageSize: 512, temporalPatchDim: 2)
// → MLXArray (n_groups, T*3=6, 512, 512) for RADIO video_embedderLicense
NVIDIA Open Model License — see the base model for full terms. Quantization, conversion, native MLX port, and runtime by Jinho Jang (eric@jangq.ai).
🦖 **Osaurus** is the open-source MLX inference server for Apple Silicon. 🌀 **JANG** is the quantization + runtime stack powering this bundle.
