wazimondo/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-GGUF
0444
Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored โ GGUF
๐ฏ GGUF Conversion of the original **AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16**
๐ฆ Files
๐ง Model Architecture
Source Model
- Creator: AEON-7
- Original: Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16
- Base Architecture: NVIDIA NemotronH (Nemotron-3-Nano)
LLM Backbone โ NemotronHForCausalLM
- Layers: 52 hybrid layers
- Hidden Size: 2,688
- MoE: 128 routed experts, 6 active per token, 1 shared expert
- Hybrid Pattern:
MEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEMEM*EMEMEMEME M= Mamba2 (SSM)E= Expert (Mixture of Experts)*= Attention- Context Length: 131,072 tokens
- Vocab Size: 131,072
Vision Encoder โ ViT-Huge (RADIO v2.5)
- Layers: 32
- Hidden Size: 1,280
- Image Size: 512ร512
- Patch Size: 16
Audio Encoder โ Parakeet
- Layers: 24
- Hidden Size: 1,024
- Sample Rate: 16kHz
๐ Usage with llama.cpp
Text-Only Inference
./llama-cli \
-m Nemotron-3-Nano-Omni-AEON-Uncensored-Q6_K.gguf \
-p "Hello, how are you?" \
-n 256Vision (Image Understanding)
./llama-llava-cli \
-m Nemotron-3-Nano-Omni-AEON-Uncensored-Q6_K.gguf \
--mmproj mmproj-Nemotron-3-Nano-Omni-BF16.gguf \
--image image.jpg \
-p "Describe this image in detail."๐ Conversion Details
๐ Credits & Attribution
This is a GGUF conversion only. All model weights originate from:
๐ก Note: This repo contains only GGUF format files converted from the original AEON-7 model. For the original PyTorch/SafeTensors weights, please visit the source model.
โ ๏ธ Disclaimer
This is an uncensored model. Use responsibly and in compliance with applicable laws and regulations.
