kabachuha/G4-MeroMero-v2-31B-Heretic-ARA-LoRA
231
G4-MeroMero-v2-31B - Heretic - LoRA
91% fewer refusals (8/100 This vs 97/100 Original) while preserving model quality (0.0785 KL divergence).
Made with the new 4bit ARA LoRA (usearalora) technique at home.
FAQ - How to use this model?
Because I don't want to spam 60+ GB uploads on huggingface, this model (natively created as lora) is distributed as a lora file. You can merge it into the original model using PEFT or, if you are using llama.cpp, you can use the convenient GGUF file (convert lora to gguf script) which is located in the "quantized" section.
llama.cpp example command:
./llama-server <your current command> --lora-scaled "/media/kabachuha/doc/G4-MeroMero-v2-31B-Heretic-ARA-LoRA.gguf:1.0"Performance
Lower refusals indicate fewer content restrictions, while lower KL divergence indicates more closeness to the original model's baseline. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections.
Abliteration parameters
Targeted components
- attn.o_proj
- mlp.down_proj
