CoolFace
Modelpublic

hiwaifu-research/WaifuGemma4-26b-a4b-v1-i1-GGUF

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
1likes5.1kdownloads
Model Card

WaifuGemma4-26b-a4b-v1 · GGUF (imatrix quants)

WaifuGemma4-26b-a4b-v1 is a role-play model: Gemma 4 26B-A4B (25.2B total, 3.8B active) post-trained with GRPO against a reward model learned from 1.2M blind votes cast by real users on the HiWaifu Arena. In blind head-to-head battles it splits votes evenly with GLM-5.1 and beats the untuned Gemma 4 26B-A4B-it 59.7% of the time. The full training story, arena results and response comparisons are on the model card.

This repository holds the weighted/imatrix quantizations. Static quants and the vision projector (mmproj) are in WaifuGemma4-26b-a4b-v1-GGUF. Compared with the static quants, the imatrix versions cut the divergence from bf16 roughly in half at 4 bits: i1-Q4KM has a mean KLD of 0.069425 against 0.150663 for static Q4KM at the same 16.8 GB, and i1-IQ4XS at 13.9 GB beats static Q4K_M at 16.8 GB.

Provided quants

QuantSize (GB)Mean KLD vs bf16Top-1 agreementPPL (RP replies)Notes
i1-Q4_K_M16.80.069491.0%5.80recommended; half the KLD of static Q4KM at the same size
i1-Q4_K_S15.50.072790.2%5.84recommended; slightly smaller
i1-IQ4_XS13.90.101589.3%6.11recommended for 16 GB cards
i1-IQ3_M12.40.145586.0%6.00matches static Q4KM quality at 12.4 GB
i1-IQ3_XXS11.30.331478.2%6.34noticeable loss; for 12 GB cards
i1-IQ2_M10.40.399777.4%6.82largest loss in the set; for 12 GB cards
imatrix0.06–––importance matrix, for making your own quants

For reference, the static quants in the other repo:

QuantSize (GB)Mean KLD vs bf16Top-1 agreementPPL (RP replies)Notes
Q8_026.90.009896.5%5.85closest to bf16; 32 GB+ VRAM or CPU RAM
Q6_K22.60.022594.3%5.76very good; the sweet spot if you have the memory
Q5_K_M19.10.058191.3%6.11good
Q4_K_M16.80.150785.9%6.12fast; the imatrix Q4KM in the i1 repo is clearly better at the same size

Running it

Any recent llama.cpp works (the GGUFs were made with build 4fea119, September 2026). The model was trained and evaluated in non-thinking mode; pass --reasoning-budget 0 so the template inserts the empty thought channel the model expects.

bash
# llama.cpp server, downloads the file straight from this repo
llama-server -hf hiwaifu-research/WaifuGemma4-26b-a4b-v1-i1-GGUF:Q4_K_M --jinja --reasoning-budget 0 -c 16384 -ngl 99

# then point SillyTavern / any OpenAI-compatible client at http://localhost:8080/v1
  • —Sampling: temperature 1.0, top-p 0.95, top-k 64 (the model's shipped defaults). 0.8–1.0 all work.
  • —Length: left alone it writes 400–500 tokens. Add Respond in no more than N tokens. to the system prompt to rein it in; it listens.
  • —System prompt: a plain character card works. Never speak or act for {{user}}. is the single most useful line.
  • —Context: trained on 8K-token conversations; the base supports 256K. Arena win rate rises with conversation depth, so long histories are fine.
  • —Vision: the static repo ships mmproj-WaifuGemma4-26b-a4b-v1-f16.gguf; download it and pass --mmproj if you want image input (or add --no-mmproj to silence the lookup). Image input was untouched by training and has not been evaluated for role-play.
  • —Content: mature content follows your system prompt. No safety tuning beyond base Gemma 4; 18+ use only.

How the quality numbers were measured

The three quality columns come from llama-perplexity --kl-divergence against the bf16 GGUF: mean KL divergence of the quant's next-token distribution from bf16's, the share of tokens where both pick the same top token, and the quant's own perplexity. The text is 20 chunks × 512 tokens of role-play replies written by this model in ten languages, held out from the imatrix calibration set.

One detail matters for anyone reproducing this. Gemma 4 instruction-tuned models, Google's included, only produce calibrated likelihoods inside a model turn; on raw text llama-perplexity reports perplexities in the tens of thousands for every Gemma 4 IT GGUF, which says nothing about the quant. We measured with the model-turn prefix <|channel>thought\n<channel|> inserted after BOS at the start of every chunk, which is what the chat template puts in front of every reply in non-thinking mode. With that prefix the bf16 model scores PPL 5.82 on the held-out replies. The patch is small: in tools/perplexity/perplexity.cpp, right after the tool replaces the first token of each chunk with BOS (in both perplexity() and kl_divergence()), overwrite the next four tokens with ids 100, 45518, 107, 101, and pass parse_special=true to common_tokenize so chat-formatted files tokenize correctly. Without it, compare quants by KL divergence only and ignore the absolute perplexity.

How these were made

  • —Converted with convert_hf_to_gguf.py --outtype bf16, then llama-quantize from the bf16 GGUF with --imatrix.
  • —The HF repo's tokenizer_config.json is in the transformers-5 format (extra_special_tokens is a list). llama.cpp pins transformers 4.57.6, which expects a dict and fails with AttributeError: 'list' object has no attribute 'keys' in _set_model_specific_special_tokens. Either install transformers>=5 in the conversion environment or change that field to {"video_token": "<|video|>"} before converting; both produce identical vocabularies. The same applies to Google's own Gemma 4 checkpoints.
  • —Quality was checked as described below before upload.
  • —The imatrix was computed on the bf16 GGUF over 1,235 chunks × 512 tokens (--parse-special) of 200 real role-play conversations from the HiWaifu Arena, formatted with the Gemma 4 chat template exactly as the model sees them at inference: system prompt, history, and the model's reply behind the non-thinking prefix. Languages: Spanish, Russian, English, Portuguese, Indonesian, Arabic, Thai, French, Ukrainian, German. The imatrix file is included so you can make your own quants.

License

Apache 2.0, as inherited from the base model.