CoolFace
Modelpublic

vcruz305/Ornith-1.0-35B-AEON-Ultimate-Uncensored-GGUF

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
17likes2.1kdownloads
Model Card

Ornith 1.0 35B AEON Ultimate Uncensored - GGUF

GGUF quantizations of AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16, produced from the BF16 source weights using importance-matrix calibration.

Two variants are provided: standard trunk (no MTP) and MTP-grafted (with Multi-Token Prediction block for speculative decoding).

Files

Standard Trunk (no MTP)

For standard autoregressive inference. Smaller files, no speculative decoding overhead.

FileQuantSizeBPWTarget GPU
ornith-aeon-35b-Q8_0.ggufQ8_035 GB~8.548GB+ (near-lossless)
ornith-aeon-35b-Q6_K.ggufQ6_K27 GB~6.632GB+ (high quality)
ornith-aeon-35b-Q5_K_M.ggufQ5KM24 GB~5.724GB (quality-first)
ornith-aeon-35b-Q4_K_M.ggufQ4KM21.2 GB~4.824GB (balanced)
ornith-aeon-35b-Q4_K_S.ggufQ4KS19.9 GB~4.624GB fallback
ornith-aeon-35b-IQ4_XS.ggufIQ4_XS18.7 GB~4.320GB (RTX 4000 Ada)
ornith-aeon-35b-Q3_K_M.ggufQ3KM16.8 GB~3.916GB GPUs
ornith-aeon-35b-Q2_K.ggufQ2_K12.9 GB~3.012GB GPUs
ornith-aeon-35b-IQ1_M.ggufIQ1_M8.2 GB~1.88GB GPUs
ornith-aeon-35b.imatrix-184 MB-Importance matrix

MTP-Grafted (with Multi-Token Prediction)

These GGUFs contain all 785 MTP tensors grafted from the base Qwen/Qwen3.5-35B-A3B model, adding a full MTP prediction block (blk.40) with 256 MoE experts. Use with --spec-type draft-mtp for speculative decoding. MTP is bundled in the GGUF -- no separate draft file needed.

FileQuantSizeBPWTarget GPU
ornith-aeon-35b-MTP-Q8_0.ggufQ8_036 GB~8.548GB+ (near-lossless)
ornith-aeon-35b-MTP-Q6_K.ggufQ6_K28 GB~6.632GB+ (high quality)
ornith-aeon-35b-MTP-Q5_K_M.ggufQ5KM24 GB~5.724GB (quality-first)
ornith-aeon-35b-MTP-Q4_K_M.ggufQ4KM21.7 GB~4.824GB (balanced)
ornith-aeon-35b-MTP-Q4_K_S.ggufQ4KS20.4 GB~4.624GB fallback
ornith-aeon-35b-MTP-IQ4_XS.ggufIQ4_XS19.2 GB~4.320-24GB
ornith-aeon-35b-MTP-Q3_K_M.ggufQ3KM17.2 GB~3.916GB GPUs
ornith-aeon-35b-MTP-Q2_K.ggufQ2_K13.3 GB~3.012GB GPUs
ornith-aeon-35b-MTP.imatrix-184 MB-Importance matrix

Why Two Variants?

The original AEON fine-tune lost its 785 MTP weight tensors. HuggingFace Transformers' AutoModelForCausalLM silently drops all mtp.* tensors during loading (_keys_to_ignore_on_load_unexpected = [r"^mtp.*"]). The mtp_num_hidden_layers: 1 in config.json is orphaned metadata from the base Qwen3.5-35B-A3B model.

The MTP-grafted variants restore all 785 MTP tensors (~488 MB BF16) by copying them from the base Qwen3.5-35B-A3B model. This works because Ornith shares identical architecture, hidden dimensions, expert count, and embedding space with its base model. Community-measured acceptance rates for grafted MTP: 58-100%.

MTP Benchmark Results

Tested: MTP-Q4_K_M on DGX Spark (GB10)

ModePromptGeneration
Standard (no MTP)47.4 t/s34.9 t/s
Embedded --spec-type draft-mtp46.2 t/s62.8 t/s

1.8x generation speedup with the bundled MTP head on a single GPU.

Embedded MTP vs separate draft model

An important distinction for MoE speculative decoding:

  • —Embedded MTP head (bundled in the GGUF, --spec-type draft-mtp): Net positive. The MTP head shares the model's KV cache and embedding space. Measured +27-80% generation speedup depending on workload and hardware.
  • —Separate draft model (-md with a standalone GGUF): Net negative for MoE. A separate draft model triggers expert-union overhead during batch verification -- more expert weight blocks must be read from VRAM, exceeding the savings. Benchmarks show -18% to -52% regression.

The MTP-grafted GGUFs in this repo use the embedded approach.

How to Run

Standard inference (recommended for most users)

bash
llama-server \
  -m ornith-aeon-35b-Q4_K_M.gguf \
  -a ornith-aeon-35b \
  --host 0.0.0.0 --port 8083 \
  -ngl 99 \
  --n-cpu-moe 3 \
  -fa on \
  -ctk q4_0 -ctv q4_0 \
  -c 32768 \
  -b 2048 -ub 768 \
  -np 1 -cb \
  -n 16384 \
  --temp 0.6 --top-k 20 --top-p 0.95 \
  --repeat-penalty 1.1 \
  --jinja \
  --reasoning-format deepseek \
  --reasoning-budget 1024

With MTP speculative decoding (faster generation)

Requires llama.cpp built from latest master (b9606+). MTP is bundled in the GGUF -- no separate draft file needed.

bash
llama-server \
  -m ornith-aeon-35b-MTP-Q4_K_M.gguf \
  -a ornith-aeon-35b \
  --spec-type draft-mtp \
  --host 0.0.0.0 --port 8083 \
  -ngl 99 \
  --n-cpu-moe 3 \
  -fa on \
  -ctk q4_0 -ctv q4_0 \
  -c 32768 \
  -b 2048 -ub 768 \
  --parallel 1 \
  --temp 0.6 --top-k 20 --top-p 0.95 \
  --repeat-penalty 1.1 \
  --jinja \
  --reasoning-format deepseek \
  --reasoning-budget 1024

RTX 4000 Ada 20GB

bash
llama-server \
  -m ornith-aeon-35b-MTP-IQ4_XS.gguf \
  -a ornith-aeon-35b \
  --spec-type draft-mtp \
  --host 0.0.0.0 --port 8083 \
  -ngl 99 \
  --n-cpu-moe 5 \
  -fa on \
  -ctk q4_0 -ctv q4_0 \
  -c 8192 \
  -b 1024 -ub 512 \
  --parallel 1 \
  --temp 0.6 --top-k 20 --top-p 0.95 \
  --repeat-penalty 1.1 \
  --jinja \
  --reasoning-format deepseek \
  --reasoning-budget 1024

High quality on 32GB+ GPU

bash
llama-server \
  -m ornith-aeon-35b-MTP-Q6_K.gguf \
  -a ornith-aeon-35b \
  --spec-type draft-mtp \
  --host 0.0.0.0 --port 8083 \
  -ngl 99 \
  -fa on \
  -ctk q8_0 -ctv q8_0 \
  -c 32768 \
  -b 2048 -ub 768 \
  --parallel 1 \
  --temp 0.6 --top-k 20 --top-p 0.95 \
  --repeat-penalty 1.1 \
  --jinja \
  --reasoning-format deepseek \
  --reasoning-budget 1024

Speculative Decoding Options

All available in llama.cpp latest master. Work on any CUDA GPU (Ada Lovelace, Ampere, Blackwell):

StrategyFlagDraft Model?Notes
MTP (embedded)--spec-type draft-mtpBundled in MTP GGUFsRequires --parallel 1. Measured +80% on Spark.
Eagle3--spec-type draft-eagle3Separate GGUFPR #18039. No pre-trained drafter for this model yet.
DFlash--spec-type draft-dflashSeparate GGUFz-lab/Qwen3.5-35B-A3B-DFlash. Best results on vLLM/SGLang at BF16/FP8.
N-gram--spec-defaultNoneZero VRAM cost. Marginal benefit on structured output.

Key Parameters

ParameterValueWhy
-ngl 99Offload all layers to GPUMaximizes speed
--spec-type draft-mtpMTP speculative decodingOnly with MTP-grafted GGUFs
--parallel 1Single slotRequired for MTP
--n-cpu-moe NOffload N MoE layers to CPUSaves ~0.8GB per layer
-fa onFlash attentionReduces KV cache memory
-ctk q4_0 -ctv q4_0Quantize KV cacheSaves ~60% KV cache VRAM
--jinjaJinja2 chat templateRequired for this model
--reasoning-format deepseekThinking modeModel uses think tags

If It OOMs

  1. 1.Increase --n-cpu-moe (trades speed for VRAM)
  2. 2.Lower context: -c 8192 or -c 4096
  3. 3.Lower batch: -b 512 -ub 256
  4. 4.Use a trunk (non-MTP) file to save ~0.5-1GB
  5. 5.Use a smaller quant

Model Details

  • —Architecture: Qwen3.5-MoE (Mixture of Experts)
  • —Total Parameters: ~35B
  • —Active Parameters per Token: ~9B (4 of 64 experts active)
  • —MTP Tensors: 785 (grafted from base model in MTP variants)
  • —Chat Template: ChatML
  • —Thinking/Reasoning: DeepSeek-style think tags
  • —Context: Up to 131K tokens (limited by available VRAM)

Quantization Details

  • —Source: BF16 safetensors from AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16
  • —MTP Graft Source: 785 mtp.* tensors from Qwen/Qwen3.5-35B-A3B (~488 MB BF16)
  • —Importance Matrix: Calibrated on diverse coding, debugging, system design, and reasoning prompts
  • —llama.cpp: Built from latest master (post-b9606, with MTP/Eagle3/DFlash support merged)
  • —Platform: DGX Spark (aarch64, CUDA 13.0)

Credits