pugant/Ornith-1.0-35B-ROCmFP4-STRIX_LEAN
Ornith-1.0-35B — ROCmFP4-STRIX_LEAN (Strix Halo / gfx1151)
Version 1.0 — 2026-08-11
TL;DR
Ornith-1.0-35B (35B params, 3B active per token, Qwen3.5-VL-MoE family) quantized to `Q4_0_ROCMFP4_STRIX_LEAN` (type 106 preset, ~4.29 BPW). Tuned for AMD Strix Halo (gfx1151 / RDNA 3.5) on the ROCmFPX fork family — we serve and benchmark these files on our lab runtime (full source: pugant/strix-nebulosa; upstream: charlie12345/ROCmFPX). Runs the full vision + text multimodal model in ~17.3 GiB.
⚠️ Critical warnings — read before downloading
- Requires a ROCmFPX fork build — see Runtime. Upstream charlie12345/ROCmFPX loads these files too (any build with the custom GGML types, e.g. built via the
kyuz0/amd-strix-halo-toolboxescontainer). The type 106 (Q4_0_ROCMFP4_STRIX_LEAN) tensor format is INVALID in stock llama.cpp — it will refuse to load. See Usage below. - Profiled for gfx1151 only (Strix Halo / Ryzen AI Max+ 395, RDNA 3.5). Not tested on other GPUs.
- FP4 here is software on RDNA 3.5 (no FP4 silicon units): the win is bandwidth / memory, not raw compute throughput. The Strix Halo ceiling on this MoE is bandwidth-bound, which is exactly where FP4 helps.
Benchmarks
Tested on Strix Halo (AMD Ryzen AI Max+ 395, 128 GB LPDDR5X). Methodology: llama-bench -ngl 999 -fa on -p 512 -n 128 -mmap 0. All benchmarks for this model were run in ROCm containers (HIP backend). Since then we also benchmarked ROCmFPX quants on the Vulkan (RADV) build of the fork — see the comparison table in Qwen3.6-35B-A3B-MTP-Q6_0_ROCMFPX; those numbers are on a different backend and not directly comparable.
vs production reference: +5.9% tok/s vs Qwen3.6-35B-A3B (66.68 vs 63), at the same 17.3 GiB footprint. See the sibling grug quant for a +12% variant (same arch family, grug fine-tune).
System configuration at bench time
Declared for reproducibility:
- Bare metal host: Bosgame BeyondMax Series, Ubuntu 24.04.4 LTS, kernel 7.0.0-28-generic
- CPU power profile:
balanced(powerprofilesctl get) — default, NOT forced to `performance`. Representative of an out-of-the-box setup. - CPU scaling driver:
amd-pstate-epp, scaling_governorperformance(amd-pstate-epp default), EPPperformance - IOMMU / iGPU power: auto (no manual tuning)
Note: tok/s above were measured on a non-tuned system (power profile balanced). Users who set powerprofilesctl set performance may see slightly higher numbers.
Quantization details
Preset Q4_0_ROCMFP4_STRIX_LEAN (GGUF file_type 106, ~4.29 bits/weight):
- Attention K/V (
blk.*.attn_qkv.weight,blk.*.attn_v.weight) →q4_0_rocmfp4(high-precision path for attention state) - Token embeddings (
token_embd.weight) →Q5_K(preserve vocab fidelity) - Expert FFN (
blk.*.ffn_*_exps.weight) →q4_0_rocmfp4_fast(max speed path; the bulk of MoE weights) - Other tensors → F32 / Q40ROCMFP4_FAST as appropriate
Reference fork: `charlie12345/ROCmFPX` commit 00d5452.
Serving runtime: see Runtime.
imatrix methodology
Precomputed by unsloth (46 chunks), redistributed here as imatrix.dat with explicit attribution. The original is at `unsloth/Ornith-1.0-35B-GGUF` (MIT).
Files
Usage
# Requires the kyuz0 Strix Halo toolbox (which builds charlie12345/ROCmFPX)
docker run --rm -p 1234:1234 --device /dev/kfd --device /dev/dri \
-v /path/to/models:/models rocmfpx-llm-service \
llama-server \
-m /models/Ornith-1.0-35B-ROCmFP4-STRIX_LEAN.gguf \
--mmproj /models/mmproj-F16.gguf \
-ngl 999 -fa on --jinja -c 32768 --host 0.0.0.0 --port 1234Notes:
- MTP not enabled at runtime. The source model includes
mtp_num_hidden_layers=1(MTP weights are present asblk.40.*), but this quant is aligned with the plain-inference MoE pipeline (no--spec-type draft-mtp). MTP weights remain in the file (~1–2 GiB extra) should a future runtime activate them. - The
--mmprojflag is required for the vision tower (multimodal). Without it, text-only still works.
How to replicate
Pipeline described in text only (no published scripts):
- Build the
docker-llm-service-convertimage fromkyuz0/amd-strix-halo-toolboxes+charlie12345/ROCmFPX(commit00d5452or later main HEAD — must containMODEL_ARCH.QWEN35MOE). - Download the BF16 GGUF (2 shards) from `unsloth/Ornith-1.0-35B-GGUF`.
- Quantize with the included
imatrix.dat:llama-quantize --imatrix imatrix.dat <bf16>.gguf <out>.gguf Q4_0_ROCMFP4_STRIX_LEAN 16.
Attribution & model tree
Qwen3.5-VL-MoE (base architecture)
└── ornith-ai/Ornith-1.0-35B (MIT)
└── this GGUF (ROCmFP4-STRIX_LEAN)- Base model: `ornith-ai/Ornith-1.0-35B` (MIT) — alias of
deepreinforce-ai/Ornith-1.0-35B - BF16 source + imatrix: `unsloth/Ornith-1.0-35B-GGUF` (MIT)
- Quantization fork: `charlie12345/ROCmFPX` (MIT)
- Container runtime: `kyuz0/amd-strix-halo-toolboxes`
License
MIT (inherited from ornith-ai/Ornith-1.0-35B and unsloth/Ornith-1.0-35B-GGUF). Derivative work: original model and its license are preserved. See `LICENSE` and `NOTICE`.
Acknowledgements
Built on the shoulders of giants:
- kyuz0/amd-strix-halo-toolboxes — Strix Halo container runtime
- charlie12345/ROCmFPX — llama.cpp fork with ROCmFP4 presets (type 106)
- unsloth — BF16 GGUF + precomputed imatrix for Ornith-1.0-35B
- ornith-ai / DeepReinforce Team — Ornith-1.0-35B
- llama.cpp community + Kawrakow (imatrix methodology)
- Hardware: Bosgame BeyondMax Series (Strix Halo bare metal host)
Limitations & community feedback
- Speed benchmark only. No perplexity / MMLU / quality eval is included in this release. The MoE structure is preserved bit-for-bit from the BF16 source except for the quantized tensor formats above; quality is expected to track standard Q4KM-class with the ROCmFP4 attention/K-V choices, but this is not measured here.
- Profiled for gfx1151 only. Not tested on other GPUs (no Navi 3 / Navi 4 / data-center MI series numbers — feel free to share yours).
- MTP present in weights but not activated at runtime (plain inference).
We invite the community — especially fellow Strix Halo owners — to test and share quality results. Open a Discussion on this repo.
Citation
@misc{ornith102026,
title = {Ornith-1.0-35B},
author = {DeepReinforce Team},
year = {2026},
url = {https://deep-reinforce.com/ornith_1_0.html}
}Disclaimer
No affiliation with AMD, Qwen, DeepReinforce, unsloth, kyuz0, or charlie12345. Provided as-is, without warranty. Users must comply with the base model license (MIT).
Runtime
Benchmark environment: one bare-metal AMD Strix Halo (Ryzen AI MAX+ 395, 128 GB) — full dated configuration and measurement policy: BARE-METAL.md
Requires a ROCmFPX fork build (custom tensor types — stock llama.cpp refuses the file). Recommended: our lab build (pugant/strix-nebulosa, main) — reasoning budget and persistent prompt cache on every model; drafter features where the model ships one: see the engine section of its README.
Vision quant served plain (MTP not enabled at runtime); prompt cache + reasoning budget still apply.
Everything here is experimental and provided as-is, at your own risk.
