malaiwah/Qwen3.8-27B-EXL3-K5K6-context
Qualify Qwen v5 body-only claims and reconcile legacy protocol summaries
Add pinned QFS size–KL plot, highlighted deep link and downloadable provenance
Add pinned QFS size–KL plot, highlighted deep link and downloadable provenance
Fix credits: RTX 5090 tester is malaiwah (self), GLM-5.2-EXL3-TR3-MTP78 is malaiwah's checkpoint; jcartu contributed the recipe/runtime that informed it
Add collection summary and fix credits (add Martin Vit, Luke Alonso, Brandon Music; correct jcartu attribution to MTP78 recipe)
Sync retrospective corrections and Final Frontier limitations
dedupe serving-profiles section (old unmarked copy removed; marked copy is canonical)
serving profiles: balanced+ANY_BITS n=3, 24/24 needles, baked image, ctx 249600
card: fidelity profile re-measured after the K5/ANY_BITS cure - prefill 1,966 -> 2,988 tok/s (+52%), decode 228.3/104.1, KLD 0.003405 (better), n=3 boots, gate 9/9
card: retract 'statistically level with official FP8' — the two KLDs come from different-sized suites and balanced's CI low end (0.005302) sits above FP8's point estimate (0.005294), so equivalence is not supportable
card: add measured serving-profile comparison + charts (measured on the -hydrated sibling, provenance noted)
charts: add fidelity-vs-quants.png
charts: add profiles-throughput.png
charts: add profiles-tradeoff.png
Address a peer review generated by our own served model: correct Three archival mirrors to Seven on four cards (all seven were already cited), and state both the build tree (21 files) and published tree (23 files) on the S16-V card with the +49,271 B accounted to README.md and DOCS-SHA256SUMS
Terminal-Bench 2.1: measured results, attribution and caveats
Correct the README row format in DOCS-SHA256SUMS so every row verifies with sha256sum -c
K6-parity: 0.001634 matches GGUF Q6_K, and the registered interval miss is published beside it
In-suite replay floor 5.83e-04, int8 embed overlay measured as an opt-in KV lever
V2 runner depth schedule measured: +38.0 % aggregate decode at C8, at the cost of 14.1 % KV tokens
Byte axis corrected: GGUF text-only file vs our multimodal payload; body-vs-body and deployed-multimodal now published
Head attribution measured: the output head is <=5.28% of divergence, so the cross-engine level gap is not decomposed
Scratch-arena overlay: measured +17,874 KV tokens (265,122 -> 282,996, +0.60 GiB) on the physical 5090, opt-in overlay not part of the qualified digest
KV-dtype sweep landed: fp8 measured default, per-token-head family measured, int4 capacity lever, nvfp4 refusal; V2-runner static-depth note
shard-0 comparator: add gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 row, paired results, third archival mirror; fix collection link
master chart: add the gittensor NVFP4 RTX5090 row; rebind DOCS-SHA256SUMS
Method hardening: publish the corpus text (corpus/text/, 941 documents + manifest) and correct the recoverability claims (audit gap G1); add the absolute-resolution paragraph beside every v5 table (gap G2); link receipts/startup-times.json beside the cold-start prose
Chat template: byte-identity to upstream, the three 400-class restrictions, prefix-reuse guidance
Link the upstream livelock issue #52520 and the boundary PR #47272
Utilisation is a per-card measurement: 0.955 is what qualified on GPU-506a575d, a second physical RTX 5090 needed 0.956, the margin is about 0.01 GiB, and a startup OOM should be answered by raising utilisation 0.001 at a time rather than dropping the window
Rebind docs checksums to the refreshed all-measurements chart
Redraw the all-measurements chart with the NVFP4 ten-shard row and the v3-vs-v5 ratio band
the sibling rebuild was scored: paired -3.755e-06 [-2.854e-05, +2.062e-05], brackets zero, so the untested-expectation wording is replaced by the measurement on all six cards
name the conversion-capable image, not an untracked path
converter determinism: a rebuild is a sibling, not the published bytes
context card: state R1's zero as a 95 % upper bound rather than an absence (0.90 per thousand builds, 1.42 per thousand gather-branch builds), per GdnGateAtConcurrency
GDN #51812 overlay recommendation on the three prefix-caching recipes, its measured absence on the context edition, and the 24 GB-class block (corrected gate clause) on the context card
Prefix caching: promote the four-module image to the release unit; enable the cache on the three 8,192-token recipes, decline it at the native 262,144 window with the measured mechanism, and correct the KV cost model to the affine per-request law
Cards: NVFP4 on v5, v3-vs-v5 scaling, capture determinism, external cross-citation, all-measurements chart, reproduce section, upstream fixes, 5090 lever sweep
Card consistency pass: resolve 16 audit findings (hardware-qualified framing, resident-vs-checkpoint bytes, whole-tree download convention, hardware-labelled throughput, figure alt text); refresh DOCS-SHA256SUMS
Bind docs checksum to normalized attributes
Hardware-qualified on a physical RTX 5090; GGUF comparison; corrected utilisation
Publish 10.48M-position KLD run, exact tail, paired MMLU-Pro matrix
Bind docs checksum to normalized attributes
Publish corrected qualification card and evidence
Chart: context edition reaches native 262,144 with int8 embeddings
Card: int4 draft embedding table is free (acceptance 56.7% vs 56.1%); MTP's cost at native context is not weights
Card: choice table reflects native 262,144
Card: native 262,144 context on a 32 GB card via an int8 embedding table (+0.000065 KLD), verified by 3/3 needle retrievals from 227,334-token prompts
Card: post-selection qualification on a source-disjoint suite - 42/42 paired wins over official FP8, ranking preserved
