SubMaroon/Gemma-4-26B-A4B-StyleTune-QK-Heretic
Gemma-4-26B-A4B-StyleTune-QK-Heretic
Gemma 4 26B-A4B with three edits: refusal directions ablated, lm_head replaced, attention routing interpolated toward a roleplay finetune.
Released for use as a starting point for further finetunes and merges.
MoE experts, router, embeddings, MLP and the vision tower are bit-for-bit identical to the abliterated body. Verified by tensor comparison.
Recipe
Alphas: 0.85 on sliding q/k, 0.65 on global q/k.
Global layers are 5, 11, 17, 23, 29. On those layers v_proj does not exist and k_proj serves as both keys and values, so a non-zero alpha on global k also injects donor content into the residual stream. Alpha there is lower than on the sliding layers for that reason, and non-zero so that queries and keys are rotated together.
StyleTune-V2 trains only lm_head; all 30 transformer layers in it are bit-identical to vanilla, checked across all 60 QK tensors before merging. That is why vanilla is used as the task-vector reference even though Pantheon-V2 was trained on top of StyleTune-V2.
Size of the QK edit
Relative Frobenius norm of the task vector against the base weights, measured before merging:
Per-row rotation between the base weights and the merged weights, in degrees, at the release alphas:
The rotation is concentrated in a small number of rows: maximum over mean runs from about 14 on globalq to about 106 on globalk.
Gemma applies RMSNorm after q_proj and k_proj, so uniform rescaling of these weights does not affect attention. Only rotation matters, which is why the angle is reported alongside the norm.
What the QK step changes
The comparison below isolates the QK step. Both arms are the same build (abliterated body, StyleTune head) and differ only in the 60 QK tensors. Effects shared with vanilla Gemma 4, and effects produced by the head transplant, are excluded.
Four prompts, two light and two dark, greedy decoding, two modes each, 2500 token budget. All generations terminated normally. Full dump included in the repo.
Prompt parsing improves. The merge more often acts on the specific situation stated in the prompt, including its participants and their actions. On a prompt about discovering that a trusted party has already signed against the character, the pre-QK build writes about the feeling of betrayal; the merge identifies the document, the mismatched handwriting and the prior meeting. On a prompt about a neighbour calling over a fence, the merge stages the neighbour's approach and the call; the pre-QK build has the neighbour present but not calling.
Lexical diversity drops slightly. Type/token ratio 0.412 before, 0.392 after, over about 2000 words per arm. Small enough that it may be noise at this sample size, but it does not move upward.
Length is unchanged. Mean prose length 1443 chars before, 1487 after. Mean word length 4.49 before, 4.52 after.
Thinking-block length moved in opposite directions on two different prompt sets and is therefore not reported.
