vcruz305/Ornith-1.0-35B-AEON-Ultimate-Uncensored-GGUF
Ornith 1.0 35B AEON Ultimate Uncensored - GGUF
GGUF quantizations of AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16, produced from the BF16 source weights using importance-matrix calibration.
Two variants are provided: standard trunk (no MTP) and MTP-grafted (with Multi-Token Prediction block for speculative decoding).
Files
Standard Trunk (no MTP)
For standard autoregressive inference. Smaller files, no speculative decoding overhead.
MTP-Grafted (with Multi-Token Prediction)
These GGUFs contain all 785 MTP tensors grafted from the base Qwen/Qwen3.5-35B-A3B model, adding a full MTP prediction block (blk.40) with 256 MoE experts. Use with --spec-type draft-mtp for speculative decoding. MTP is bundled in the GGUF -- no separate draft file needed.
Why Two Variants?
The original AEON fine-tune lost its 785 MTP weight tensors. HuggingFace Transformers' AutoModelForCausalLM silently drops all mtp.* tensors during loading (_keys_to_ignore_on_load_unexpected = [r"^mtp.*"]). The mtp_num_hidden_layers: 1 in config.json is orphaned metadata from the base Qwen3.5-35B-A3B model.
The MTP-grafted variants restore all 785 MTP tensors (~488 MB BF16) by copying them from the base Qwen3.5-35B-A3B model. This works because Ornith shares identical architecture, hidden dimensions, expert count, and embedding space with its base model. Community-measured acceptance rates for grafted MTP: 58-100%.
MTP Benchmark Results
Tested: MTP-Q4_K_M on DGX Spark (GB10)
1.8x generation speedup with the bundled MTP head on a single GPU.
Embedded MTP vs separate draft model
An important distinction for MoE speculative decoding:
- Embedded MTP head (bundled in the GGUF,
--spec-type draft-mtp): Net positive. The MTP head shares the model's KV cache and embedding space. Measured +27-80% generation speedup depending on workload and hardware. - Separate draft model (
-mdwith a standalone GGUF): Net negative for MoE. A separate draft model triggers expert-union overhead during batch verification -- more expert weight blocks must be read from VRAM, exceeding the savings. Benchmarks show -18% to -52% regression.
The MTP-grafted GGUFs in this repo use the embedded approach.
How to Run
Standard inference (recommended for most users)
llama-server \
-m ornith-aeon-35b-Q4_K_M.gguf \
-a ornith-aeon-35b \
--host 0.0.0.0 --port 8083 \
-ngl 99 \
--n-cpu-moe 3 \
-fa on \
-ctk q4_0 -ctv q4_0 \
-c 32768 \
-b 2048 -ub 768 \
-np 1 -cb \
-n 16384 \
--temp 0.6 --top-k 20 --top-p 0.95 \
--repeat-penalty 1.1 \
--jinja \
--reasoning-format deepseek \
--reasoning-budget 1024With MTP speculative decoding (faster generation)
Requires llama.cpp built from latest master (b9606+). MTP is bundled in the GGUF -- no separate draft file needed.
llama-server \
-m ornith-aeon-35b-MTP-Q4_K_M.gguf \
-a ornith-aeon-35b \
--spec-type draft-mtp \
--host 0.0.0.0 --port 8083 \
-ngl 99 \
--n-cpu-moe 3 \
-fa on \
-ctk q4_0 -ctv q4_0 \
-c 32768 \
-b 2048 -ub 768 \
--parallel 1 \
--temp 0.6 --top-k 20 --top-p 0.95 \
--repeat-penalty 1.1 \
--jinja \
--reasoning-format deepseek \
--reasoning-budget 1024RTX 4000 Ada 20GB
llama-server \
-m ornith-aeon-35b-MTP-IQ4_XS.gguf \
-a ornith-aeon-35b \
--spec-type draft-mtp \
--host 0.0.0.0 --port 8083 \
-ngl 99 \
--n-cpu-moe 5 \
-fa on \
-ctk q4_0 -ctv q4_0 \
-c 8192 \
-b 1024 -ub 512 \
--parallel 1 \
--temp 0.6 --top-k 20 --top-p 0.95 \
--repeat-penalty 1.1 \
--jinja \
--reasoning-format deepseek \
--reasoning-budget 1024High quality on 32GB+ GPU
llama-server \
-m ornith-aeon-35b-MTP-Q6_K.gguf \
-a ornith-aeon-35b \
--spec-type draft-mtp \
--host 0.0.0.0 --port 8083 \
-ngl 99 \
-fa on \
-ctk q8_0 -ctv q8_0 \
-c 32768 \
-b 2048 -ub 768 \
--parallel 1 \
--temp 0.6 --top-k 20 --top-p 0.95 \
--repeat-penalty 1.1 \
--jinja \
--reasoning-format deepseek \
--reasoning-budget 1024Speculative Decoding Options
All available in llama.cpp latest master. Work on any CUDA GPU (Ada Lovelace, Ampere, Blackwell):
Key Parameters
If It OOMs
- Increase
--n-cpu-moe(trades speed for VRAM) - Lower context:
-c 8192or-c 4096 - Lower batch:
-b 512 -ub 256 - Use a trunk (non-MTP) file to save ~0.5-1GB
- Use a smaller quant
Model Details
- Architecture: Qwen3.5-MoE (Mixture of Experts)
- Total Parameters: ~35B
- Active Parameters per Token: ~9B (4 of 64 experts active)
- MTP Tensors: 785 (grafted from base model in MTP variants)
- Chat Template: ChatML
- Thinking/Reasoning: DeepSeek-style think tags
- Context: Up to 131K tokens (limited by available VRAM)
Quantization Details
- Source: BF16 safetensors from AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16
- MTP Graft Source: 785 mtp.* tensors from Qwen/Qwen3.5-35B-A3B (~488 MB BF16)
- Importance Matrix: Calibrated on diverse coding, debugging, system design, and reasoning prompts
- llama.cpp: Built from latest master (post-b9606, with MTP/Eagle3/DFlash support merged)
- Platform: DGX Spark (aarch64, CUDA 13.0)
Credits
- Base model by AEON-7
- MTP weights from Qwen/Qwen3.5-35B-A3B
- Quantization by vcruz305
