veldierin/Artemis-31B-v1q-GGUF
<!-- ARTEMIS model card — Artemis palette. silver #9fc2e6 · amber #e0b15c · deep navy #05080f · ax- class scheme. --> <div class="ax-page">
<div class="ax-hero"><img src="artemis-hero-4.png" alt="ARTEMIS — 31B · Gemma 4 · Pure Roleplay" /></div>
<h1>ARTEMIS-31B-v1q</h1> <div class="ax-chips"> <span class="ax-chip">31B GEMMA 4</span> <span class="ax-chip">GEMMA4 ARCH</span> <span class="ax-chip">I1 IMATRIX</span> </div>
<p class="ax-lead">A pure roleplay finetune of <code>google/gemma-4-31B-it</code>. Trained by <b>Drummer</b> on his own core roleplay datasets, hosted here by <b>BeaverAI</b> as a quantized GGUF for local engines.</p>
<div class="ax-grid"> <div class="ax-card ax-left"> <h3>What this is</h3> <p>Artemis is a single-purpose build: no generalist benchmark chasing, no "agentic" garnish — the weight updates all go toward roleplay. It inherits the Gemma 4 base (arch <code>gemma4</code>, 833 tensors) and keeps the full upstream chat template, so it drops straight into any Gemma-family launcher.</p> </div> <div class="ax-card ax-right"> <h3>What it isn't</h3> <p>No public benchmark numbers exist for this finetune and none are claimed here. Judge it on the actual generation: in-character consistency, agency, and writing quality in long roleplay scenes.</p> </div> </div>
<h2>Specs</h2> <table class="ax-table"> <thead><tr><th>Property</th><th>Value</th></tr></thead> <tbody> <tr><td>Base</td><td><code>google/gemma-4-31B-it</code></td></tr> <tr><td>Finetune</td><td>v1q — Drummer core roleplay datasets</td></tr> <tr><td>Architecture</td><td><code>gemma4</code> · 833 tensors</td></tr> <tr><td>Chat template</td><td>Preserved from base (intact through quant)</td></tr> <tr><td>Purpose</td><td>Pure roleplay</td></tr> </tbody> </table>
<h2>Available precisions</h2> <table class="ax-table"> <thead><tr><th>File</th><th>Quant</th><th>Size</th><th>Notes</th></tr></thead> <tbody> <tr><td><code>Artemis-31B-v1q-i1-bt-IQ4XS.gguf</code></td><td>IQ4XS (imatrix · bartowski)</td><td>15.98 GiB</td><td>Dense ternary-float4 — the workhorse build, scotoma-calibrated</td></tr> <tr><td><code>Artemis-31B-v1q-i1-mr-IQ4XS.gguf</code></td><td>IQ4XS (imatrix · mradermacher)</td><td>15.59 GiB</td><td>Same precision, mradermacher base-model matrix</td></tr> <tr><td><code>Artemis-31B-v1q-i1-bt-IQ3M.gguf</code></td><td>IQ3M (imatrix · bartowski)</td><td>14.09 GiB</td><td>IQ3S base + bartowski selective upcast recipe (31×Q6K, 67×Q4K, 50×Q5K)</td></tr> <tr><td><code>Artemis-31B-v1q-i1-bt-IQ3XXS.gguf</code></td><td>IQ3XXS (imatrix · bartowski)</td><td>11.92 GiB</td><td>IQ3XXS base + 91 bartowski selective upcasts (Q6K attnq, IQ3S attnoutput, Q5K token_embd)</td></tr> </tbody> </table>
<h2>Speculative Decoding</h2> <p class="ax-lead">Supports SPEC-MTP speculative decoding via a separate draft head. Load with: <code>--spec-type draft-mtp -md ../gemma-4-31B-it-Q8_0-MTP.gguf</code> in llama.cpp or by selecting the draft head in your launcher. Works with any precision listed above. Recommended: <code>--spec-draft-n-max 2</code> to balance acceptance rate and quality; values >2 may degrade long-range coherence.</p>
<h2>Archival</h2> <p class="ax-lead">The following build uses the <b>v1p</b> checkpoint (previous iteration). It is included for users with existing v1p pipelines and is not the current recommended build.</p> <ul class="ax-list" style="list-style: none; padding: 0; margin: 0.5rem 0;"> <li><code>Artemis-31B-v1p-i1-bt-IQ3XXS.gguf</code> — v1p at IQ3XXS (11.92 GiB)</li> </ul>
<h2>Credits</h2> <table class="ax-table"> <thead><tr><th>Role</th><th>Credit</th></tr></thead> <tbody> <tr><td>Finetune & core datasets</td><td><b>Drummer</b></td></tr> <tr><td>Hosting & distribution</td><td><b>BeaverAI</b></td></tr> <tr><td>Quantization</td><td>imatrix-assisted GGUF — IQ4XS, IQ3M + IQ3_XXS (bartowski + mradermacher matrices)</td></tr> </tbody> </table>
<p class="ax-foot">Trained by Drummer · Hosted by BeaverAI · GGUF for llama.cpp-family engines.</p>
</div>
