JamieBradfield/qwen3.8-9b-hermes-fc-balanced
Qwen3.8-9B Hermes FC — Balanced
A QLoRA continue-tune of the v2 real-traces model (qwen3.8-9b-hermes-fc-real-traces), trained on a balanced mix of real Hermes agent trajectories plus big-model teacher data. Part of an iterative recipe series (v1 → v2/real-traces → v26/balanced) teaching a 9B model to drive the Hermes agent toolset.
- Base model: Empero/Qwen3.8-9B (Apache-2.0)
- Warm start: v2 merged weights (
qwen3.8-9b-hermes-fc-real-traces), QLoRA adapter on top (envelope + tool tokens already baked into v2's weights) - Training data: 1,543 rows in ShareGPT format:
- 979 v2 rows (933 real Hermes agent trajectories + 46 gap-fill), lambda rows dropped
- SWE-rebench windows (Qwen3-Coder-480B OpenHands trajectories remapped onto real Hermes tools:
execute_bash→terminal,str_replace_editor→read_file/write_file/patch) - APIGen-MT-5k call-units (GPT-4o / DeepSeek-V3)
- When2Call PREF rows (Mixtral-8x22B)
- Method: QLoRA 4-bit, rank 16 / alpha 16, dropout 0, batch 1 x grad-accum 8 (effective 8), lr 2e-4, warmup 0.1, MAX_SEQ 6144, 1 epoch = 193 steps (final loss 0.0552)
- Vocabulary: 248077 → 248079 (2 added tool tokens:
<|tool_call|>,<|tool_response|>) - Weights: BF16 full merge (12 shards, 18.4 GB, 9.20B params) plus a ROCmFPX quantized GGUF (see companion repo)
- Architecture note: merge preserves the 15-key MTP (multi-token prediction) head from the base —
Qwen3_5ForConditionalGenerationwithtext_config.vocab_size = 248079. Visual tower dropped (text-only).
Addendum 2026-09-03 — evaluation-methodology correction
Read this before you use the numbers below.
After this checkpoint was published, we tested the base model (Empero/Qwen3.8-9B-Distill, no fine-tuning) through the live Hermes interface: the OpenAI-native tools parameter path that Hermes actually uses. The base model calls tools correctly on its own.
The harness that certified this checkpoint did not use that interface. It baked tool schemas into the system text, never passed tools, and required raw <tool_call> XML in the response. That prompt shape matches neither the base model's own chat template nor the live Hermes request path. Every gate score in the v26-to-v29 lineage measured this synthetic interface.
The statement in this README that "the base distill contributes zero tool-calling" is incorrect. It measured the synthetic interface.
Corrected comparison (native tools path, same 4-tool battery, greedy):
What the fine-tune lineage actually changed: todo-first planning on multi-step tasks (tier-2 name=todo 3/10 base to 8/10 v28 to 10/10 armB2). That is real, and it is the one capability the lineage added. The costs: restraint regressions (tier-3 over-trigger 5/10 to 9/10; tier-4 tool-drift on trivia 0/10 to 4/10) and, in armB2, over-suppression (tier-1 10/20).
Status: the fine-tune program is frozen as of 2026-09-03. The base model is the author's runtime default. This checkpoint stays published as the strongest measured demonstration of SFT-taught todo-first planning in the lineage, and as the artifact that exposed the harness error. The Evaluation section below describes behavior on the synthetic interface; do not read it as native-interface competence.
Reproduction: scripts/native_battery.py + scripts/todo_probes.jsonl
scripts/no_drift_probes.jsonlproduce the corrected table (any GGUF,CTX=65536to match the serve config).
- Base model: Empero/Qwen3.8-9B (Apache-2.0)
- Warm start: v27 merged weights (
qwen3.8-9b-hermes-fc-todo), QLoRA adapter on top (envelope + tool tokens already baked in) - Training data: 218 trajectories (ShareGPT format, 551 todo calls), generated by GLM-5.3-Flash via the HuggingFace router on 112 synthetic tasks mined from the author's agent corpus — not the author's own traces:
- 190 teacher trajectories (85% verified by sandbox execution; failures with recovery injected for 15.8% of rows)
- 61 rows carry distilled short-chain-of-thought; 159 envelope-only
- 12 synthetic web-search tasks (first web signal)
- Full dataset published separately (see
hermes-fc-v28-data); probes were kept virgin (zero overlap with generation or training) - Method: QLoRA 4-bit, rank 16 / alpha 16, dropout 0, batch 1 x grad-accum 8 (effective 8), lr 2e-4, warmup 0.1, MAX_SEQ 6144, 8 epochs = 214 steps (final loss 0.02085, avg train loss 0.09405)
- Vocabulary: 248077 → 248079 (2 added tool tokens:
<|tool_call|>,<|tool_response|>) - Weights: BF16 full merge (13 shards, 18.9 GB, 9.20B params) plus a ROCmFPX quantized GGUF (see companion repo)
- Architecture note: merge preserves the 15-key MTP (multi-token prediction) head from the base —
Qwen3_5ForConditionalGenerationwithtext_config.vocab_size = 248079. Visual tower dropped (text-only).
Evaluation (held-out probes, greedy)
45 real ambiguous trajectories (seeded held-out split), each demanding one of a specific tool family. Same seed/limit for all models.
v26 fires more (43/45 vs 41/45) and formats more exactly (27/45) than v2, with clarify intact (2/2). Cost: −3 namematch / −2 argsok — the model sometimes picks the wrong tool under semantic ambiguity (e.g. search_files instead of tool_describe on a docs probe). Selection precision is the target of the next recipe iteration.
The base distill contributes zero tool-calling (0/45); the entire capability comes from the fine-tune.
Intended use
Research artifact for experimenters working on tool-call behavior in 9B-class models — not a product. The tool_call XML envelope and tool schemas match the Hermes agent runtime this was trained on.
What this repo contains
- The BF16 merged weights (this repo)
- All scripts that produced the model and dataset (
scripts/)
Reproducibility
All scripts that produced this model are in `scripts/`:
The scripts carry the author's machine paths (F:/models/..., C:/AI/...) and run on this rig's ROCm stack (Unsloth 2026.8.18, torch 2.11 ROCm, RX 7700 XT); adjust paths for your environment.
Convert/quantize: llama-rocmfpx fork, convert_hf_to_gguf.py --outtype bf16 → llama-quantize Q4_0_ROCMFP4_FAST (see GGUF repo).
Quantized version
Q4_0_ROCMFP4_FAST GGUF (ROCmFPX format, AMD RDNA3 kernels) is published in the companion repo. ROCmFPX quants target AMD ROCm inference; for portable use, convert from the BF16 merge here.
Training data note
The training data itself is not published. The rows derive from the author's own agent sessions, and some rows contain private strings (hostnames, session identifiers). The scripts that build the dataset are published; the data is not. The weights do not contain those strings.
Acknowledgements
Base model and license inherited from Empero/Qwen3.8-9B (Apache-2.0). Training data derived from the author's own Hermes agent sessions plus publicly-licensed teacher sources (SWE-rebench, APIGen-MT-5k, When2Call).
