CoolFace
Modelpublic

wazimondo/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-GGUF

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
0likes444downloads
Model Card

Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored โ€” GGUF

๐ŸŽฏ GGUF Conversion of the original **AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16**

๐Ÿ“ฆ Files

FileFormatSizeDescription
Nemotron-3-Nano-Omni-AEON-Uncensored-BF16.ggufBF16~59GBFull precision LLM weights
Nemotron-3-Nano-Omni-AEON-Uncensored-Q6_K.ggufQ6_K~32GBQuantized LLM weights (recommended)
mmproj-Nemotron-3-Nano-Omni-BF16.ggufBF16~1.5GBVision encoder projection

๐Ÿง  Model Architecture

Source Model

LLM Backbone โ€” NemotronHForCausalLM

  • โ€”Layers: 52 hybrid layers
  • โ€”Hidden Size: 2,688
  • โ€”MoE: 128 routed experts, 6 active per token, 1 shared expert
  • โ€”Hybrid Pattern: MEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEM*EMEMEMEM*EMEMEMEME
  • โ€”M = Mamba2 (SSM)
  • โ€”E = Expert (Mixture of Experts)
  • โ€”* = Attention
  • โ€”Context Length: 131,072 tokens
  • โ€”Vocab Size: 131,072

Vision Encoder โ€” ViT-Huge (RADIO v2.5)

  • โ€”Layers: 32
  • โ€”Hidden Size: 1,280
  • โ€”Image Size: 512ร—512
  • โ€”Patch Size: 16

Audio Encoder โ€” Parakeet

  • โ€”Layers: 24
  • โ€”Hidden Size: 1,024
  • โ€”Sample Rate: 16kHz

๐Ÿš€ Usage with llama.cpp

Text-Only Inference

bash
./llama-cli \
  -m Nemotron-3-Nano-Omni-AEON-Uncensored-Q6_K.gguf \
  -p "Hello, how are you?" \
  -n 256

Vision (Image Understanding)

bash
./llama-llava-cli \
  -m Nemotron-3-Nano-Omni-AEON-Uncensored-Q6_K.gguf \
  --mmproj mmproj-Nemotron-3-Nano-Omni-BF16.gguf \
  --image image.jpg \
  -p "Describe this image in detail."

๐Ÿ“ Conversion Details

StepToolDetails
LLM ExtractionPython safetensorsExtracted language_model.* from omni checkpoint
GGUF Conversionconvert_hf_to_gguf.pyllama.cpp โ€” NemotronHForCausalLM โ†’ BF16 GGUF
Quantizationllama-quantizeBF16 โ†’ Q6_K
mmprojconvert_hf_to_gguf.py --mmprojNemotronNanoV2VLModel โ€” ViT + projector โ†’ BF16

๐Ÿ™ Credits & Attribution

This is a GGUF conversion only. All model weights originate from:

RoleEntity
๐Ÿง  Original Model**AEON-7**
๐Ÿ“ ArchitectureNVIDIA NemotronH
๐Ÿ”ง GGUF Conversion@hotdogs
๐Ÿ› ๏ธ Toolsllama.cpp by Georgi Gerganov & contributors
๐Ÿ’ก Note: This repo contains only GGUF format files converted from the original AEON-7 model. For the original PyTorch/SafeTensors weights, please visit the source model.

โš ๏ธ Disclaimer

This is an uncensored model. Use responsibly and in compliance with applicable laws and regulations.