JamieBradfield/qwen3.8-9b-hermes-fc-tooluse-GGUF
Qwen3.8-9B Hermes FC — Tooluse (GGUF)
ROCmFPX-quantized version of `JamieBradfield/qwen3.8-9b-hermes-fc-tooluse` (BF16 merge in the parent repo).
Quant details
Convert/quantize: llama-rocmfpx fork, convert_hf_to_gguf.py --outtype bf16 → llama-quantize Q4_0_ROCMFP4_FAST.
ROCmFPX quants target AMD ROCm inference (RX 7700 XT in the author's rig, 12 GB VRAM, served at 64k context with q8_0/turbo3 KV — the measured sweet spot: ~5.5x decode speed vs 245k context). For portable use, convert from the BF16 merge in the parent repo instead.
Addendum 2026-09-03 — evaluation-methodology correction
The evaluation claims on the original card for this repo came from a synthetic harness that baked tool schemas into the system text and never passed the OpenAI tools parameter. Through the native tools path (the interface live Hermes uses), the base model — no fine-tuning — fires tier-1 at 18/20 with zero tier-4 drift, and this model's tier-2 todo-first advantage shrinks to 8/10 vs the base's 3/10 (see the full corrected table in the parent repo card). The fine-tune program is frozen as of 2026-09-03; the base model is the author's runtime default.
Note on the MTP head
The quant preserves the 15-tensor MTP (multi-token prediction) head from the base under the mtp.* tensor prefix (442 tensors total in this GGUF), usable with --spec-type draft-mtp on supporting builds.
