hiwaifu-research/WaifuGemma4-26b-a4b-v1-i1-GGUF
WaifuGemma4-26b-a4b-v1 · GGUF (imatrix quants)
WaifuGemma4-26b-a4b-v1 is a role-play model: Gemma 4 26B-A4B (25.2B total, 3.8B active) post-trained with GRPO against a reward model learned from 1.2M blind votes cast by real users on the HiWaifu Arena. In blind head-to-head battles it splits votes evenly with GLM-5.1 and beats the untuned Gemma 4 26B-A4B-it 59.7% of the time. The full training story, arena results and response comparisons are on the model card.
This repository holds the weighted/imatrix quantizations. Static quants and the vision projector (mmproj) are in WaifuGemma4-26b-a4b-v1-GGUF. Compared with the static quants, the imatrix versions cut the divergence from bf16 roughly in half at 4 bits: i1-Q4KM has a mean KLD of 0.069425 against 0.150663 for static Q4KM at the same 16.8 GB, and i1-IQ4XS at 13.9 GB beats static Q4K_M at 16.8 GB.
Provided quants
For reference, the static quants in the other repo:
Running it
Any recent llama.cpp works (the GGUFs were made with build 4fea119, September 2026). The model was trained and evaluated in non-thinking mode; pass --reasoning-budget 0 so the template inserts the empty thought channel the model expects.
# llama.cpp server, downloads the file straight from this repo
llama-server -hf hiwaifu-research/WaifuGemma4-26b-a4b-v1-i1-GGUF:Q4_K_M --jinja --reasoning-budget 0 -c 16384 -ngl 99
# then point SillyTavern / any OpenAI-compatible client at http://localhost:8080/v1- Sampling: temperature 1.0, top-p 0.95, top-k 64 (the model's shipped defaults). 0.8–1.0 all work.
- Length: left alone it writes 400–500 tokens. Add
Respond in no more than N tokens.to the system prompt to rein it in; it listens. - System prompt: a plain character card works.
Never speak or act for {{user}}.is the single most useful line. - Context: trained on 8K-token conversations; the base supports 256K. Arena win rate rises with conversation depth, so long histories are fine.
- Vision: the static repo ships
mmproj-WaifuGemma4-26b-a4b-v1-f16.gguf; download it and pass--mmprojif you want image input (or add--no-mmprojto silence the lookup). Image input was untouched by training and has not been evaluated for role-play. - Content: mature content follows your system prompt. No safety tuning beyond base Gemma 4; 18+ use only.
How the quality numbers were measured
The three quality columns come from llama-perplexity --kl-divergence against the bf16 GGUF: mean KL divergence of the quant's next-token distribution from bf16's, the share of tokens where both pick the same top token, and the quant's own perplexity. The text is 20 chunks × 512 tokens of role-play replies written by this model in ten languages, held out from the imatrix calibration set.
One detail matters for anyone reproducing this. Gemma 4 instruction-tuned models, Google's included, only produce calibrated likelihoods inside a model turn; on raw text llama-perplexity reports perplexities in the tens of thousands for every Gemma 4 IT GGUF, which says nothing about the quant. We measured with the model-turn prefix <|channel>thought\n<channel|> inserted after BOS at the start of every chunk, which is what the chat template puts in front of every reply in non-thinking mode. With that prefix the bf16 model scores PPL 5.82 on the held-out replies. The patch is small: in tools/perplexity/perplexity.cpp, right after the tool replaces the first token of each chunk with BOS (in both perplexity() and kl_divergence()), overwrite the next four tokens with ids 100, 45518, 107, 101, and pass parse_special=true to common_tokenize so chat-formatted files tokenize correctly. Without it, compare quants by KL divergence only and ignore the absolute perplexity.
How these were made
- Converted with
convert_hf_to_gguf.py --outtype bf16, thenllama-quantizefrom the bf16 GGUF with--imatrix. - The HF repo's
tokenizer_config.jsonis in the transformers-5 format (extra_special_tokensis a list). llama.cpp pins transformers 4.57.6, which expects a dict and fails withAttributeError: 'list' object has no attribute 'keys'in_set_model_specific_special_tokens. Either installtransformers>=5in the conversion environment or change that field to{"video_token": "<|video|>"}before converting; both produce identical vocabularies. The same applies to Google's own Gemma 4 checkpoints. - Quality was checked as described below before upload.
- The imatrix was computed on the bf16 GGUF over 1,235 chunks × 512 tokens (
--parse-special) of 200 real role-play conversations from the HiWaifu Arena, formatted with the Gemma 4 chat template exactly as the model sees them at inference: system prompt, history, and the model's reply behind the non-thinking prefix. Languages: Spanish, Russian, English, Portuguese, Indonesian, Arabic, Thai, French, Ukrainian, German. The imatrix file is included so you can make your own quants.
License
Apache 2.0, as inherited from the base model.
