batiai/Nemotron-3-Nano-30B-A3B-GGUF
Nemotron 3 Nano 30B-A3B GGUF — Quantized by BatiAI
<p align="center"> <a href="https://flow.bati.ai"><img src="https://img.shields.io/badge/BatiFlow-macOS%20AI%20Automation-blue?style=for-the-badge&logo=apple" alt="BatiFlow"></a> <a href="https://ollama.com/batiai/nemotron3-nano"><img src="https://img.shields.io/badge/Ollama-batiai%2Fnemotron3--nano-green?style=for-the-badge" alt="Ollama"></a> </p>
Quantizations of NVIDIA Nemotron 3 Nano 30B-A3B (NemotronH MoE, hybrid Mamba+Attention) for on-device AI on Mac. Built and verified by BatiAI for BatiFlow.
Why Nemotron 3 Nano?
- 30B params, only 3B active per token — A3B MoE architecture
- Hybrid Mamba + Attention — long context with linear scaling
- Reasoning + agentic — built for tool use, structured outputs
- Apache-spirit NVIDIA Open Model License — commercial-friendly
- Runs on a 32GB Mac with IQ3/IQ4
Quick Start
ollama pull batiai/nemotron3-nano:iq4Available Quantizations
RAM Requirements
Model Comparison — Which BatiAI Model for Your Mac?
Why Nemotron-H Architecture?
NemotronH is NVIDIA's hybrid architecture combining Mamba state-space layers with standard attention:
- Linear scaling on long context (Mamba) + accuracy at short context (Attention)
- A3B MoE: 128 experts, 8 active per token
- 49 layers with hybrid override pattern
- Trained on reasoning, code, and agentic data
Why BatiAI Quantization?
Technical Details
- Original Model: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
- Architecture: NemotronH MoE (Mamba + Attention hybrid, 30B-A3B)
- License: NVIDIA Open Model License
- Quantized with: llama.cpp —
imatrix --chunks 200calibrated - Quantized by: BatiAI
About BatiFlow
BatiFlow — free, on-device AI automation for Mac. 5MB app, 100% local, unlimited. 57+ built-in tools for calendar, notes, reminders, files, email, browser, messaging.
License
Quantized from nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16. License: NVIDIA Open Model License.
Benchmarks
<!-- BENCH-START -->
<!-- BENCH-END -->
