AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
docs: recommend Qwen3.8 MIXED + GH recipe card
docs: redirect to Qwen3.8 NVFP4-MIXED successor
Recipe: gpu-util 0.6-0.7 on DGX Spark unified memory (>~0.8 thrashes the shared pool); discrete VRAM unchanged
add AEON Qwen cover art
AEON Qwen cover art
docs: Spark DFlash recipe -> n=10 + mamba float32 (drop block-size 256); MTP n=3 path unchanged
Add MLX variants + Apple Silicon hardware routing to Variants table
tags: expand to maximally-searchable set (+38 tags, union with existing)
docs(quickstart): comprehensive top quickstart (pull container+model+drafter) + AGENTS.md v0.23.0
docs(quickstart): unify on aeon-vllm-ultimate:latest + validated serve flags (gpu-util<=0.88)
docs(perf): v0.23.0 benchmarks on aeon-vllm-ultimate:latest + charts
Container association -> aeon-vllm-ultimate:latest (DFlash n=12); BF16 path unchanged
docs: add AEON vLLM Ultimate (v0.22.1 + PR #44389 NVFP4 KV) container reference (sibling of validated MTP-XS)
Tip jar: single left-aligned QR column (fix narrow-viewport clipping)
README: document MTP head graft (2026-05-01); credit @tcclaviger (#6)
Update index.json for grafted MTP head (15 new entries)
Graft MTP head (15 tensors) from Qwen/Qwen3.6-27B base — credit @tcclaviger (#6)
Add tip jar block (BTC/ETH/SOL/XMR with QR codes)
Restore base Qwen 3.6 pre_tokenizer (\p{M}) + trim_offsets:false; keep AEON's audio/TTS added tokens. Fixes multilingual tokenization regression flagged in BF16 discussion #5; matches what the base was trained with so combining-mark languages (Hindi, Arabic, Vietnamese, etc.) tokenize correctly. English/code unchanged.
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Upload README.md with huggingface_hub
Variants table: add Multimodal-NVFP4-MTP + Text-NVFP4-MTP entries with experimental disclaimer
Add Variants table + Precision/quantization config block; rename to -BF16 (per community feedback HF discussion #4)
Add prominent GitHub-repo callout under title (deployment guide, AGENTS.md, benchmarks live there)
Rewrite intro with build-investment narrative, fix refusal denominator to 0/100, add vLLM serving defaults for 80/96 GB GPUs
Elaborate model card: full stats, capability writeup, arbitration clause
Upload folder using huggingface_hub
initial commit
