CoolFace
Modelpublic

hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290

sourceHugging Faceapache-2.0updated 5mo agoView on Hugging Face
0likes27downloads
Model Card

๐Ÿง  Gemma 4 E4B Ultra Uncensored Heretic โ€” Unsloth QLoRA (r=16)

Model ID: hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290

A lightweight LoRA adapter (rank 16) fine-tuned with QLoRA on llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic using Unsloth. Trained on reasoning traces distilled from Claude Opus for 3 epochs (1,335 steps) โ€” just 73 MB (F16 GGUF) / 162 MB (safetensors).


๐Ÿ“Š Training Summary

MetricValue
Base Modelllmfan46/gemma-4-E4B-it-ultra-uncensored-heretic
Datasetlordx64/reasoning-distill-claude-opus-4-7-max
Training TypeQLoRA (load_in_4bit: true)
Epochs3.0 (1,335 steps)
Train Loss13.32 โ†’ 2.22 (โ†“ 83%)
Final Step Loss1.28 (step 1,335)
Eval Loss2.88
Learning Rate2e-4 โ†’ cosine decay โ†’ 1.5e-7
Total Tokens Seen1,608,612
Training Time~7.2 hours
HardwareNVIDIA RTX 4060 Ti 16GB
CUDA / Driver13.0 / 580.126.09

๐Ÿ› ๏ธ LoRA Configuration

ParameterValue
Rank (`r`)16
Alpha16 (lora_alpha / r = 1.0)
Dropout0.0
Target Modulesq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Biasnone
PEFT Version0.18.1

Additional Training Settings

ParameterValue
Batch Size1 (effective 18 with gradient accumulation)
Max Seq Length512
Optimizeradamw_bnb_8bit
LR Schedulerlinear
Warmup Steps5
Weight Decay0.001
Random Seed3407
Sequence Packingโœ… enabled
Gradient Checkpointingunsloth

๐Ÿ“ˆ Loss Curve

Step     0:  13.32  โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ
Step   300:   2.13  โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–
Step   600:   1.64  โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ
Step   900:   1.59  โ–ˆโ–ˆโ–ˆโ–ˆโ–Š
Step  1200:   1.61  โ–ˆโ–ˆโ–ˆโ–ˆโ–‰
Step  1335:   1.28  โ–ˆโ–ˆโ–ˆโ–‰  โ† FINAL

Training converged smoothly from initial loss ~13.3 down to 2.22 (average). The final training step achieved 1.28 loss. Eval loss at 2.88 suggests moderate overfitting common with small LoRA adapters โ€” expected and acceptable for the adapter size (73 MB).


๐Ÿš€ How to Use

Option 1: PEFT (PyTorch)

python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_model = "llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic"
lora_path  = "hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290"

model = AutoModelForCausalLM.from_pretrained(
    base_model,
    torch_dtype="auto",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(base_model)

model.load_adapter(lora_path, adapter_name="lora")
model.set_active_adapter("lora")

messages = [{"role": "user", "content": "Explain the theory of relativity simply."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Option 2: Unsloth (Recommended โ€” 2ร— faster, uses less VRAM)

python
from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290",
    max_seq_length=2048,
    load_in_4bit=True,  # or False for BF16
)
FastLanguageModel.for_inference(model)

messages = [{"role": "user", "content": "Write a poem about AI in Thai."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to("cuda")

output = model.generate(**inputs, max_new_tokens=512, temperature=0.7)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Option 3: GGUF (llama.cpp)

Download gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf (73 MB) from the gguf/ directory. Then use with your existing base model GGUF:

bash
# Serve with llama.cpp LoRA support (llama-server with --lora)
llama-server \
  -m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
  --lora gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf \
  --lora-scaled gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf 1.0

๐Ÿ“ฆ Model Files

FileFormatSize
adapter_model.safetensorsPEFT safetensors162 MB
adapter_config.jsonPEFT config1.3 KB
gguf/gemma-4-E4B-uncensored-heretic-lora-r16.f16.ggufGGUF LoRA (F16)73.4 MB
tokenizer.jsonTokenizer31 MB
trainer_state.jsonTraining log394 KB

โš ๏ธ Limitations & Bias

  • โ€”LoRA Adapter only โ€” you need the base model `llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic` loaded separately (no merged weights included).
  • โ€”Eval gap โ€” train loss 1.28 vs eval loss 2.88 indicates some overfitting on the Claude reasoning dataset.
  • โ€”Limited context โ€” trained with max_seq_length=512 and sequence packing. Performance on very long reasoning chains may degrade.
  • โ€”Uncensored โ€” the base model has minimal alignment filtering, so outputs may be more creative/unfiltered than standard models.
  • โ€”No formal benchmarks โ€” MMLU, GSM8K, etc. not evaluated. Loss-based convergence suggests improved reasoning over the base model.
  • โ€”Single GPU โ€” trained on one RTX 4060 Ti 16GB with batch_size=1. Larger-scale generalization may vary.

๐Ÿ“š Citation

If you use this model in research or production, please credit:

bibtex
@misc{gemma4-e4b-heretic-lora-2025,
  author = {UKA (Hermes Agent)},
  title = {Gemma 4 E4B Ultra Uncensored Heretic โ€” Unsloth QLoRA Fine-tuned
           on Claude Reasoning Distill},
  year = {2025},
  publisher = {Hugging Face},
  howpublished = {\\url{https://huggingface.co/hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290}},
  note = {Trained with Unsloth on RTX 4060 Ti. Base model by llmfan46.}
}

๐Ÿ”— Links


๐Ÿ“ Changelog

DateEvent
2026-05-05 09:33Training started on llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic
2026-05-05 16:46Training completed โ€” checkpoint-1335 (3 epochs, 1,608,612 tokens)
2026-05-05 22:43GGUF LoRA exported โ€” 73.4 MB F16

Made with ๐Ÿ’œ by UKA ยท Powered by Unsloth & NVIDIA RTX 4060 Ti