CoolFace
Modelpublic

SC117/occamy-1.0-abliterated-FIT-GGUF

sourceHugging Faceapache-2.0updated 2d agoView on Hugging Face
6likes2.3kdownloads
Model Card

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 24px;"> <div style="background: #0f172a; border-radius: 20px; padding: 40px 32px 30px 32px; text-align: center; position: relative; overflow: hidden;"> <div style="position: absolute; top: -40px; right: -40px; width: 160px; height: 160px; background: rgba(139,92,246,0.18); border-radius: 50%;"></div> <div style="position: absolute; bottom: -50px; left: -30px; width: 140px; height: 140px; background: rgba(59,130,246,0.15); border-radius: 50%;"></div> <div style="position: absolute; top: 30%; right: 12%; width: 56px; height: 56px; background: rgba(59,130,246,0.22); border-radius: 50%;"></div> <div style="display: inline-flex; flex-wrap: wrap; justify-content: center; gap: 8px; margin-bottom: 18px; position: relative; z-index: 1;"><span style="background: #3b82f6; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">FIT-GGUF v0.4.0</span><span style="background: #8b5cf6; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">GATE-VERIFIED TIERS</span><span style="background: #10b981; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">5 FIT + 4 APEX-I- + BF16</span><span style="background: #ef4444; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">ABLITERATED (INPUT-GATED)</span><span style="background: #f59e0b; color: #0f172a; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">MEASURED KL + SAME-TOP</span><span style="background: #0ea5e9; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">VISION mmproj</span><span style="background: #1e293b; color: #94a3b8; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">APACHE-2.0</span></div> <h1 style="margin: 0 0 10px 0; font-size: 34px; font-weight: 800; color: #f8fafc; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">Occamy-1.0-abliterated · FIT-GGUF</h1> <p style="margin: 0 0 20px 0; font-size: 15px; color: #cbd5e1; position: relative; z-index: 1;">Five fidelity tiers of a 256K-context 35B MoE with vision — the smallest GGUF that provably meets each KL gate, re-verified on its own bytes. Plus an independent APEX-I- family and the BF16 reference, all measured by the same protocol.</p> <div style="position: relative; z-index: 1; margin: 0 auto; width: 72%; max-width: 540px;"> <div style="height: 8px; border-radius: 8px; background: #334155; position: relative; overflow: visible;"> <div style="height: 8px; width: 100%; border-radius: 8px; background: linear-gradient(90deg, #3b82f6, #8b5cf6);"></div> <div style="position: absolute; top: 50%; left: 0%; transform: translate(-50%, -50%); width: 18px; height: 18px; border-radius: 50%; background: #ffffff; box-shadow: 0 0 0 4px rgba(59,130,246,0.35);"></div> <div style="position: absolute; top: 50%; left: 100%; transform: translate(-50%, -50%); width: 18px; height: 18px; border-radius: 50%; background: #ffffff; box-shadow: 0 0 0 4px rgba(139,92,246,0.35);"></div> </div> <div style="display: flex; justify-content: space-between; margin-top: 8px; font-size: 10px; color: #94a3b8; font-weight: 600;"><span>10.94 GiB · MINI</span><span style="color: #c4b5fd;">verified minimum at each fidelity</span><span>64.61 GiB · BF16</span></div> </div> <p style="margin: 18px 0 0 0; font-size: 13px; position: relative; z-index: 1;"><span style="color: #94a3b8;">English</span><span style="color: #475569;"> · </span><a href="https://huggingface.co/SC117/occamy-1.0-abliterated-FIT-GGUF/blob/main/README.zh-CN.md" style="color: #93c5fd; text-decoration: none; font-weight: 600;">简体中文 📖</a></p> </div> </div>

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 18px; margin-bottom: 30px;">

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧭</span> About FIT-GGUF — the tool behind these files</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;">Every FIT tier here was planned, executed and verified by <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">FIT-GGUF</a>, an open-source, deterministic tensor-level planning layer on top of llama.cpp quantization. Standard GGUF quantization asks you to pick one of a handful of presets; FIT-GGUF instead asks <i>what quality do you want</i>, then finds and <b>verifies</b> the smallest GGUF that demonstrably meets it.</p><p style="margin: 0 0 12px 0; padding: 10px 14px; background: #f5f3ff; border-left: 4px solid #7c3aed; border-radius: 6px; color: #4c1d95; font-weight: 600;">Traditional GGUF gives you presets. FIT gives you a fidelity contract: macro KL ≤ tier anchor, measured against this model's own BF16 on five fixed domains.</p><table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Deterministic size prediction &amp; byte-exact delivery</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #166534; font-weight: 700;">✅ Validated (G2 gate, delta = 0)</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Universally optimal tensor allocation</td><td style="padding: 6px 10px; color: #92400e; font-weight: 700;">⚠️ Not established — FIT claims verified fidelity contracts, not a universal quality optimum</td></tr></tbody></table><p style="margin: 12px 0 0 0;">Method, pre-registered research record and the <code>fit</code> CLI are all open source: <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">github.com/Scorp1o117/FIT-GGUF</a></p></div></div>

<div style="border: 1px solid #fca5a5; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #dc2626 0%, #ea580c 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>⚠️</span> Safety notice / 安全提示</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 8px 0;">The source model is an <b>abliterated, refusal-removed</b> model with no meaningful built-in guardrails, and may comply with harmful, illegal or unsafe requests. Use it only where you can provide appropriate moderation, access control and legal review. Do not deploy it to end users without your own safety layer.</p><p style="margin: 0; color: #64748b;">源模型经过拒答方向消融,不具备可靠的内置安全护栏;请仅在合法、受控、具备审核与滥用防护的环境中使用,使用者自行承担部署责任。</p></div></div>

<div style="border: 1px solid #c4b5fd; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #7c3aed 0%, #a855f7 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧬</span> Abliteration — input-gated, documented</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;">The refusal direction was ablated locally with the <b>abliterix</b> pipeline using <b>input-gated ablation</b>. A classic rank-1 edit is <code>ΔW = −w · v (vᵀW)</code> — the same output direction <code>v</code> for <i>every</i> input, harmless ones included, which is where the KL bill comes from. The gated form replaces the input-side vector with a per-expert direction <code>g</code> chosen to maximise refusal suppression per unit of perturbation, <code>g ∝ Sh⁻¹ μr</code>: the harmful-prompt mean <b>whitened by the harmless second-moment matrix</b>, so the edit lands where harmless traffic does not live. 2,160 expert deltas were applied across 40 layers, and the export is re-derived from the original checkpoint plus those deltas and compared byte-for-byte.</p><table style="width: 100%; border-collapse: collapse; font-size: 12.5px;"><thead><tr style="background: rgba(124,58,237,0.05);"><th style="padding: 7px 10px; border-bottom: 2px solid #7c3aed; text-align: left; color: #6d28d9;">Metric</th><th style="padding: 7px 10px; border-bottom: 2px solid #7c3aed; text-align: left; color: #6d28d9;">Delivered</th></tr></thead><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">Refusals (100 held-out harmful prompts)</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700; color: #166534;">10 (99 before ablation)</td></tr><tr><td style="padding: 6px 10px; font-weight: 600;">Ablation perturbation, 3-token full-distribution KL</td><td style="padding: 6px 10px; font-weight: 700;">0.0844</td></tr><tr><td style="padding: 6px 10px; font-weight: 600;">Previous best on this model</td><td style="padding: 6px 10px; color: #64748b;">0.2196 — the gated direction cut it by <b>62%</b></td></tr></tbody></table><p style="margin: 10px 0 0 0; padding: 9px 13px; background: #fffbeb; border-left: 4px solid #f59e0b; border-radius: 6px; color: #92400e;"><b>Two different KLs — do not mix them.</b> The <b>0.0844</b> above is the <i>abliteration's</i> perturbation of the full-precision model. Every other KL number on this page is <i>quantization</i> divergence against this repository's BF16 file. They are measured on different objects and are not comparable.</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> Everything in this repository</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 12px 0;">Five FIT tiers, four APEX-I- tiers and the BF16 reference — <b>every file measured by the same evaluator, on the same five slices, against the same reference logits</b>, so the columns are directly comparable across families.</p><table style="width: 100%; border-collapse: collapse; font-size: 12.5px;"><thead><tr style="background: rgba(37,99,235,0.05);"><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Family</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">File</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">GiB</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Macro KL ↓</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Same-top ↑</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Role / gate</th></tr></thead><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700; color: #1d4ed8;">FIT REFERENCE</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-family: monospace; font-size: 11.5px;">…-FIT-REFERENCE-24G-Q5KM.gguf</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700;">23.55</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700;">0.0195</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">95.40%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #7c3aed;">KL ≤ 0.02 · near-lossless</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #1d4ed8;">FIT QUALITY</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-FIT-QUALITY-17G-IQ4XS.gguf</td><td style="padding: 6px 10px; font-weight: 700;">17.50</td><td style="padding: 6px 10px; font-weight: 700;">0.0464</td><td style="padding: 6px 10px;">92.26%</td><td style="padding: 6px 10px; color: #7c3aed;">KL ≤ 0.05 · best all-round</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #1d4ed8;">FIT BALANCED</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-FIT-BALANCED-14G-IQ3S.gguf</td><td style="padding: 6px 10px; font-weight: 700;">13.80</td><td style="padding: 6px 10px; font-weight: 700;">0.0959</td><td style="padding: 6px 10px;">88.83%</td><td style="padding: 6px 10px; color: #7c3aed;">KL ≤ 0.10 · 16 GB cards</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #1d4ed8;">FIT COMPACT</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-FIT-COMPACT-13G-IQ3XXS.gguf</td><td style="padding: 6px 10px; font-weight: 700;">12.76</td><td style="padding: 6px 10px; font-weight: 700;">0.1358</td><td style="padding: 6px 10px;">86.15%</td><td style="padding: 6px 10px; color: #7c3aed;">KL ≤ 0.15 · tighter VRAM</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #1d4ed8;">FIT MINI</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-FIT-MINI-11G-IQ2S.gguf</td><td style="padding: 6px 10px; font-weight: 700;">10.94</td><td style="padding: 6px 10px; font-weight: 700;">0.1992</td><td style="padding: 6px 10px;">83.18%</td><td style="padding: 6px 10px; color: #7c3aed;">KL ≤ 0.20 · smallest usable</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #047857;">APEX-I-balanced</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-APEX-I-BALANCED-Q6K.gguf</td><td style="padding: 6px 10px; font-weight: 700;">23.59</td><td style="padding: 6px 10px; font-weight: 700;">0.0197</td><td style="padding: 6px 10px;">95.04%</td><td style="padding: 6px 10px; color: #047857;">Q6K base</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #047857;">APEX-I-quality</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-APEX-I-QUALITY-Q6K.gguf</td><td style="padding: 6px 10px; font-weight: 700;">21.25</td><td style="padding: 6px 10px; font-weight: 700;">0.0240</td><td style="padding: 6px 10px;">94.46%</td><td style="padding: 6px 10px; color: #047857;">Q6K base</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #047857;">APEX-I-compact</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-APEX-I-COMPACT-Q4KM.gguf</td><td style="padding: 6px 10px; font-weight: 700;">15.40</td><td style="padding: 6px 10px; font-weight: 700;">0.0700</td><td style="padding: 6px 10px;">90.34%</td><td style="padding: 6px 10px; color: #047857;">Q4KM base</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #047857;">APEX-I-mini</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-APEX-I-MINI-Q3KM.gguf</td><td style="padding: 6px 10px; font-weight: 700;">12.54</td><td style="padding: 6px 10px; font-weight: 700;">0.1654</td><td style="padding: 6px 10px;">84.72%</td><td style="padding: 6px 10px; color: #047857;">Q3KM base</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #64748b;">BF16</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-BF16.gguf</td><td style="padding: 6px 10px; font-weight: 700;">64.61</td><td style="padding: 6px 10px; font-weight: 700;">0</td><td style="padding: 6px 10px;">100%</td><td style="padding: 6px 10px; color: #64748b;">The reference every number above is measured against</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #64748b;">imatrix</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">…-BF16-imatrix.gguf</td><td style="padding: 6px 10px; font-weight: 700;">0.18</td><td style="padding: 6px 10px; color: #94a3b8;">—</td><td style="padding: 6px 10px; color: #94a3b8;">—</td><td style="padding: 6px 10px; color: #64748b;">500×512 chunks, 510 entries — re-quantize this model yourself</td></tr><tr><td style="padding: 6px 10px; font-weight: 700; color: #64748b;">mmproj</td><td style="padding: 6px 10px; font-family: monospace; font-size: 11.5px;">mmproj-…-BF16.gguf</td><td style="padding: 6px 10px; font-weight: 700;">0.84</td><td style="padding: 6px 10px; color: #94a3b8;">—</td><td style="padding: 6px 10px; color: #94a3b8;">—</td><td style="padding: 6px 10px; color: #64748b;">Vision projector, BF16 — pair it with any tier above</td></tr></tbody></table><p style="margin: 12px 0 0 0; padding: 9px 13px; background: #fffbeb; border-left: 4px solid #f59e0b; border-radius: 6px; color: #92400e; font-size: 12.5px;"><b>File size ≠ RAM/VRAM usage.</b> KV cache and compute buffers are separate — and 30 of this model's 40 layers are linear-attention/DeltaNet, so only 10 expand the cache. Pick a sane <code>-c</code>.</p><p style="margin: 10px 0 0 0; font-size: 12.5px; color: #475569;"><b>The FIT tiers are 78.55 GiB</b> where the smallest passing standard preset in each would be <b>89.89 GiB</b> — <b>11.34 GiB (12.6%) smaller</b>.<p style="margin: 12px 0 0 0; font-size: 12.5px; color: #475569;">The FIT and APEX-I- tiers are <b>different points on the same quality curve</b> — different sizes, different KL. They are not ranked against each other here; the chart below shows where each one sits.</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📈</span> Measured quality</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><img src="https://huggingface.co/SC117/occamy-1.0-abliterated-FIT-GGUF/resolve/main/results/fit-vs-apex-vs-native-en.png" alt="Macro KL against main GGUF size: native presets, the five FIT tiers, and the four APEX-I- tiers" style="width: 100%; border-radius: 8px; border: 1px solid #e2e8f0;"><p style="margin: 8px 0 0 0; font-size: 12px; color: #64748b;">llama.cpp b10666 · <code>c=512 b=512</code> · five domains · vs the BF16 in this repo · <a href="https://huggingface.co/SC117/occamy-1.0-abliterated-FIT-GGUF/resolve/main/results/fit-vs-apex-vs-native-zh.png" style="color: #1d4ed8; text-decoration: none;">中文大图</a></p><p style="margin: 10px 0 0 0;">The native ladder is the reference curve; the FIT tiers sit below and to the left of it at every gate. The dashed APEX-I- line is the independent recipe family from the table above, and it is the reason the 0.02 anchor is worth reading carefully: at that gate the two families are level, while everywhere below it FIT is well clear.</p><p style="margin: 10px 0 0 0; padding: 9px 13px; background: #f0f9ff; border-left: 4px solid #2563eb; border-radius: 6px; color: #1e3a8a; font-size: 12.5px;"><b>Read the curve, not just the tiers.</b> The gates are deliberately capped at KL 0.02–0.20, so the big native presets are <i>cleaner</i> than every tier — Q6K at 26.56 GiB is 10% cleaner than REFERENCE. If you want maximum fidelity rather than minimum size, take Q6K or the BF16 file.</p><p style="margin: 0; font-size: 12px;"><a href="https://huggingface.co/SC117/occamy-1.0-abliterated-FIT-GGUF/resolve/main/results/fit-vs-apex-vs-native-en.png" style="color: #1d4ed8; text-decoration: none;">Full-size EN</a> · <a href="https://huggingface.co/SC117/occamy-1.0-abliterated-FIT-GGUF/resolve/main/results/fit-vs-apex-vs-native-zh.png" style="color: #1d4ed8; text-decoration: none;">中文大图</a> · <a href="https://huggingface.co/SC117/occamy-1.0-abliterated-FIT-GGUF/blob/main/results/emit-report.json" style="color: #1d4ed8; text-decoration: none;">raw per-domain JSON</a></p></div></div>

<div style="border: 1px solid #fca5a5; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #dc2626 0%, #ea580c 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🚀</span> Run it</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0; padding: 9px 13px; background: #f0fdf4; border-left: 4px solid #16a34a; border-radius: 6px; color: #14532d;"><b>This is plain <code>qwen35moe</code> architecture.</b> Unlike models that need a patched llama.cpp, these files load in any reasonably recent upstream build — no PR branch, no custom fork. The measurements here used <b>b10666</b>.</p><p style="margin: 0 0 8px 0; font-weight: bold; color: #1e293b;">Text only</p><p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">./llama-server -m occamy-1.0-abliterated-FIT-BALANCED-14G-IQ3_S.gguf -ngl 99 -c 32768

./llama-cli -m occamy-1.0-abliterated-FIT-MINI-11G-IQ2S.gguf -ngl 99 -c 8192 -p "Hello" -n 256</p><p style="margin: 12px 0 8px 0; font-weight: bold; color: #1e293b;">With vision</p><p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">./llama-server -m occamy-1.0-abliterated-FIT-QUALITY-17G-IQ4XS.gguf \ --mmproj mmproj-occamy-1.0-abliterated-BF16.gguf -ngl 99 -c 32768</p><p style="margin: 10px 0 0 0;">The chat template is embedded in every file — including the tool-calling and reasoning-content handling — so LM Studio, KoboldCpp, Jan and friends load them without extra configuration. Native context is 262,144 tokens, but KV cache grows with it: start at <code>-c 32768</code> and work up.</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔬</span> Evaluation protocol &amp; honest scope</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><table style="width: 100%; border-collapse: collapse; font-size: 12.5px;"><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Runtime</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">llama.cpp <b>b10666</b> · Linux x8664 · ROCm 10.0 · AMD Ryzen AI MAX+ 395 (gfx1151)</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Command shape</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;"><code>llama-perplexity -ngl 99 -t 16 -c 512 -b 512 --kl-divergence …</code></td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Reference</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">This model's own BF16 logits (the BF16 GGUF shipped here)</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Domains</td><td style="padding: 6px 10px; color: #4b5563;">wikitest · wikivalid · Chinese · code · agentchat (five fixed 64 KiB slices, macro mean)</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Calibration</td><td style="padding: 6px 10px; color: #4b5563;">12-point standard preset ladder plus gap probes; tier search by size bisection; guard profile scope <code>exactmodel</code></td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Weights binding</td><td style="padding: 6px 10px; color: #4b5563;"><code>sourceweights_sha256 = b77f1175…f78856</code> — the registry entry is keyed to these exact weights</td></tr></tbody></table><p style="margin: 10px 0 0 0;"><b>Tier verification:</b> a FIT tier ships only if macro KL ≤ its anchor. Each shipped file was re-evaluated <i>on its own bytes</i> after quantization and reproduced its search-time KL exactly (G2 delta +0 on all five — the delivered byte count equals the re-finalized prediction to the byte).</p><p style="margin: 8px 0 0 0;"><b>Allocator scope:</b> the stock <code>balanced</code> v0.3 policy with precision floors applied where the window's candidate set cannot reach a tensor, using this model's own importance matrix. No model-specific refine profile. What is claimed is deterministic size planning plus measured verification of these specific artifacts — not a universally optimal allocation.</p><p style="margin: 8px 0 0 0;"><b>The FIT tiers are deliberately not policy-uniform,</b> and the ledger records which planning policy produced each one. A tier's product is the smallest artifact that reaches its anchor <i>and can be rebuilt from the bundle</i>; that artifact does not care which policy planned it. Policy governs where the <i>search</i> looks next, and the bracket is filtered by it — filtering the <i>selection</i> by policy would have thrown away smaller passing artifacts already in the ledger.</p><p style="margin: 8px 0 0 0;"><b>Sizes in the file names are measured bytes, not budgets.</b> The tool refuses to emit an artifact whose size it cannot predict to the byte.</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧩</span> Included — and not included</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 6px 0;">✅ 5 gate-verified FIT tiers · ✅ <b>4 APEX-I- tiers</b> for the cross-check · ✅ <b>BF16 reference GGUF</b> (the exact weights every measurement above is taken against) · ✅ the <b>importance matrix</b> used for every quantization · ✅ a <b>vision projector</b> converted from the base model's visual tower · ✅ chat template embedded in every file · ✅ checksums (<code>SHA256SUMS.txt</code>) · ✅ labelled quality curves and the raw per-domain JSON (<code>results/</code>)</p><p style="margin: 0 0 6px 0;">❌ no 2-bit-and-below classes — measured collapse on this model, excluded by design · ❌ no MTP head · ❌ no GGUF above Q6_K for the text weights; take the BF16 file if you need the exact weights</p><p style="margin: 0;">The abliteration was performed locally from <a href="https://huggingface.co/Accio-Lab/occamy-1.0" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">Accio-Lab/occamy-1.0</a>; this repository contributes the abliterated weights' quantization plans, the artifacts, the vision projector and the measurements.</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔍</span> Verify &amp; reproduce</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">sha256sum -c SHA256SUMS.txt</p><p style="margin: 10px 0 0 0;">The evaluation slices are public in the <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">FIT-GGUF repository</a>; the calibration ladder, gap probes, search audit, the per-tier recipes and the guard profile for this release are retained in the FIT-GGUF experiment record <a href="https://github.com/Scorp1o117/FIT-GGUF/tree/main/experiments/2026-09-22-occamy-1p0-abliterated-4tier" style="color: #1d4ed8; text-decoration: none; font-weight: 700;"><code>experiments/2026-09-22-occamy-1p0-abliterated-4tier/</code></a>. The five-domain <code>.kld</code> reference logits are not stored there — they regenerate from the BF16 GGUF shipped here and are checked against <code>reference-manifest.json</code>, which pins each domain's hash and byte count. Because the importance-matrix path is embedded in the GGUF metadata, the <i>byte count</i> of a re-quantization depends on where you put the imatrix file — the weights do not. Exact-size behaviour is scoped to the recorded source metadata and the recorded llama.cpp build; changing the converter, runtime, source layout or metadata requires revalidation.</p></div></div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📄</span> License &amp; credits</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 8px 0;">Apache-2.0, inherited from the base model — follow the upstream license and model-card requirements.</p><p style="margin: 0;"><a href="https://huggingface.co/Accio-Lab" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">Accio-Lab</a> — the Occamy-1.0 model and its technical report · <a href="https://github.com/ggml-org/llama.cpp" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">llama.cpp</a> — quantization and the KL/perplexity evaluator · <b>abliterix</b> — the local input-gated ablation pipeline · <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">FIT-GGUF</a> — verified size-exact quantization. FIT-GGUF is an independent project, not affiliated with Accio-Lab or llama.cpp.</p></div></div> </div>

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; text-align: center; margin-bottom: 30px;"><a href="https://github.com/Scorp1o117/FIT-GGUF" style="display: inline-block; background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); color: white; font-weight: 700; font-size: 14px; padding: 12px 28px; border-radius: 24px; text-decoration: none;">⭐ FIT-GGUF on GitHub — the tool, the method, the full research record</a><p style="margin: 14px 0 0 0; font-size: 13px;"><span style="color: #86868b;">English</span><span style="color: #cbd5e1;"> · </span><a href="https://huggingface.co/SC117/occamy-1.0-abliterated-FIT-GGUF/blob/main/README.zh-CN.md" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">简体中文</a></p></div>