hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290
๐ง Gemma 4 E4B Ultra Uncensored Heretic โ Unsloth QLoRA (r=16)
Model ID: hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290
A lightweight LoRA adapter (rank 16) fine-tuned with QLoRA on llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic using Unsloth. Trained on reasoning traces distilled from Claude Opus for 3 epochs (1,335 steps) โ just 73 MB (F16 GGUF) / 162 MB (safetensors).
๐ Training Summary
๐ ๏ธ LoRA Configuration
Additional Training Settings
๐ Loss Curve
Step 0: 13.32 โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Step 300: 2.13 โโโโโโโ
Step 600: 1.64 โโโโโ
Step 900: 1.59 โโโโโ
Step 1200: 1.61 โโโโโ
Step 1335: 1.28 โโโโ โ FINALTraining converged smoothly from initial loss ~13.3 down to 2.22 (average). The final training step achieved 1.28 loss. Eval loss at 2.88 suggests moderate overfitting common with small LoRA adapters โ expected and acceptable for the adapter size (73 MB).
๐ How to Use
Option 1: PEFT (PyTorch)
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model = "llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic"
lora_path = "hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290"
model = AutoModelForCausalLM.from_pretrained(
base_model,
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(base_model)
model.load_adapter(lora_path, adapter_name="lora")
model.set_active_adapter("lora")
messages = [{"role": "user", "content": "Explain the theory of relativity simply."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))Option 2: Unsloth (Recommended โ 2ร faster, uses less VRAM)
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290",
max_seq_length=2048,
load_in_4bit=True, # or False for BF16
)
FastLanguageModel.for_inference(model)
messages = [{"role": "user", "content": "Write a poem about AI in Thai."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to("cuda")
output = model.generate(**inputs, max_new_tokens=512, temperature=0.7)
print(tokenizer.decode(output[0], skip_special_tokens=True))Option 3: GGUF (llama.cpp)
Download gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf (73 MB) from the gguf/ directory. Then use with your existing base model GGUF:
# Serve with llama.cpp LoRA support (llama-server with --lora)
llama-server \
-m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
--lora gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf \
--lora-scaled gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf 1.0๐ฆ Model Files
โ ๏ธ Limitations & Bias
- LoRA Adapter only โ you need the base model `llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic` loaded separately (no merged weights included).
- Eval gap โ train loss 1.28 vs eval loss 2.88 indicates some overfitting on the Claude reasoning dataset.
- Limited context โ trained with
max_seq_length=512and sequence packing. Performance on very long reasoning chains may degrade. - Uncensored โ the base model has minimal alignment filtering, so outputs may be more creative/unfiltered than standard models.
- No formal benchmarks โ MMLU, GSM8K, etc. not evaluated. Loss-based convergence suggests improved reasoning over the base model.
- Single GPU โ trained on one RTX 4060 Ti 16GB with
batch_size=1. Larger-scale generalization may vary.
๐ Citation
If you use this model in research or production, please credit:
@misc{gemma4-e4b-heretic-lora-2025,
author = {UKA (Hermes Agent)},
title = {Gemma 4 E4B Ultra Uncensored Heretic โ Unsloth QLoRA Fine-tuned
on Claude Reasoning Distill},
year = {2025},
publisher = {Hugging Face},
howpublished = {\\url{https://huggingface.co/hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290}},
note = {Trained with Unsloth on RTX 4060 Ti. Base model by llmfan46.}
}๐ Links
๐ Changelog
Made with ๐ by UKA ยท Powered by Unsloth & NVIDIA RTX 4060 Ti
