CoolFace
Modelpublic

darthcrawl/Lilith-31B-v1.0-GGUF

sourceHugging Facegemmaupdated 2mo agoView on Hugging Face
2likes864downloads
Model Card

<p align="center"><img src="lilith.png" width="480" alt="Lilith"/></p>

Lilith-31B-v1.0 (GGUF)

Versatile uncensored roleplay / creative-writing model on Gemma-4-31B. Drives any character card (SillyTavern / pluma), built to be lively and varied rather than flat or repetitive. darthcrawl. Explicit-capable.

  • —Base: coder3101/gemma-4-31B-it-heretic (decensored Gemma-4-31B-it)
  • —Method: QLoRA r32/alpha64 all-linear, loss masked to model turns, 1 epoch eval-driven. Release = the ~0.39-epoch checkpoint (eval 2.11 vs base 7.03), picked over the fully-trained one to keep prose lively (loss != liveliness).
  • —Data: ~30M tokens, 8 curated sources (human forum RP, AO3, curated public RP, curated synth), ~40/60 NSFW/SFW.
  • —Chat template: Gemma-4 (<|turn>user ... <turn|> / <|turn>model ...).
  • —Sampling tip: DRY + XTC + modest repetition penalty for max variety.

Siblings: Lilith-31B-v1.0 (bf16) | -LoRA | -GGUF | -MLX-4bit/6bit/8bit

Quants

FileSizeNotes
Q3KM15.3 GBtight VRAM
Q4KM18.7 GBbalanced
Q5KM21.8 GB
Q6_K25.2 GBfits 32GB w/ --cache-type-k/v q8_0
Q8_032.6 GBnear-lossless
IQ4_XS16.7 GBimatrix, best quality-per-GB at 4-bit
IQ3_M14.4 GBimatrix, tightest usable VRAM

IQ quants use a register-calibrated imatrix (3MB of held-out corpus prose, imat_register.dat in this repo). Serve with llama.cpp --jinja.