rajveer43/gemma-4-E4B-medical-legal-finance-qa
Gemma 4 E4B — Medical · Legal · Finance Domain QA
Fine-tuned version of `google/gemma-4-E4B-it` across three professional domains — Medical, Legal, and Finance — using QLoRA (4-bit NF4) with Optuna-tuned hyperparameters, trained on Kaggle T4 GPU.
Trained by rajveer43 · Apache 2.0 license
What this model does
This model answers domain-specific professional questions across three verticals. Domain context is passed via the system prompt — the same adapter handles all three.
Model details
Training pipeline
medical_meadow_medqa + legal_case_qa + finance-alpaca
│
▼
Format to Gemma 4 chat template
(system prompt encodes domain context)
│
▼
google/gemma-4-E4B-it
+ BitsAndBytes 4-bit NF4
+ LoRA adapters (r=16, RSLoRA, ~1% params)
│
▼
Optuna HPO — 8 trials × 100 steps
TPE sampler + MedianPruner
Search: lr, batch, grad_accum, warmup
│
▼
SFTTrainer — 3 epochs, cosine LR
best checkpoint saved per epoch
│
▼
GGUF Q4_K_M → llama.cppStack: unsloth · trl · peft · bitsandbytes · accelerate · optuna · datasets
Quickstart
Option A — Unsloth (fastest)
from unsloth import FastModel
import torch
model, tokenizer = FastModel.from_pretrained(
model_name = "rajveer43/gemma-4-E4B-medical-legal-finance-qa",
load_in_4bit = True,
)
FastModel.for_inference(model)
SYSTEM_PROMPTS = {
"medical": (
"You are a highly knowledgeable medical AI assistant. "
"Answer clinical and biomedical questions accurately. "
"Always advise consulting a licensed physician for personal medical decisions."
),
"legal": (
"You are an expert legal AI assistant. "
"Provide accurate information about laws, precedents, and legal concepts. "
"Always clarify responses are not a substitute for licensed legal counsel."
),
"finance": (
"You are a knowledgeable financial AI assistant. "
"Answer questions about markets, instruments, and economic concepts accurately. "
"Remind users this is not personalized financial advice."
),
}
def ask(question: str, domain: str = "medical", max_new_tokens: int = 256) -> str:
messages = [
{"role": "system", "content": SYSTEM_PROMPTS[domain]},
{"role": "user", "content": question},
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
# FIX: Bypass the Unsloth-patched multimodal processor entirely.
# For text-only inference, call the underlying fast tokenizer directly
# so the image processor pipeline is never triggered.
underlying_tokenizer = tokenizer.tokenizer # unwrap: Gemma4Processor -> PreTrainedTokenizerFast
inputs = underlying_tokenizer(
prompt,
return_tensors="pt",
padding=True,
).to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
temperature=0.1,
top_p=0.9,
do_sample=True,
)
input_len = inputs["input_ids"].shape[-1]
return underlying_tokenizer.decode(output[0][input_len:], skip_special_tokens=True)
# Medical
print(ask("What is the first-line treatment for Type 2 diabetes?", domain="medical"))
# Legal
print(ask("What constitutes a breach of contract under common law?", domain="legal"))
# Finance
print(ask("What is the difference between ROE and ROA?", domain="finance"))Option B — PEFT + Transformers
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch
bnb_config = BitsAndBytesConfig(
load_in_4bit = True,
bnb_4bit_quant_type = "nf4",
bnb_4bit_compute_dtype = torch.bfloat16,
bnb_4bit_use_double_quant = True,
)
base = AutoModelForCausalLM.from_pretrained(
"google/gemma-4-E4B-it",
quantization_config = bnb_config,
device_map = "auto",
)
model = PeftModel.from_pretrained(base, "rajveer43/gemma-4-E4B-medical-legal-finance-qa")
tokenizer = AutoTokenizer.from_pretrained("rajveer43/gemma-4-E4B-medical-legal-finance-qa")Option C — GGUF on M4 Mac (llama.cpp + Metal)
# Install
brew install llama.cpp
# Download GGUF
huggingface-cli download \
rajveer43/gemma-4-E4B-medical-legal-finance-qa \
--include "*.gguf" \
--local-dir ~/models/gemma4-domain-qa
# Interactive chat
llama-cli \
-m ~/models/gemma4-domain-qa/model-unsloth.Q4_K_M.gguf \
-ngl 99 \
--chat-template gemma \
-p "You are a knowledgeable medical AI assistant." \
-i
# OpenAI-compatible server
llama-server \
-m ~/models/gemma4-domain-qa/model-unsloth.Q4_K_M.gguf \
-ngl 99 \
--host 0.0.0.0 --port 8080Expected on M4 Mac: 35–50 tokens/sec · ~4–5 GB unified memory
Domain coverage
Medical
- First-line treatment protocols and clinical guidelines
- Drug mechanism of action and pharmacology
- Diagnostic criteria (sepsis, diabetes, hypertension, etc.)
- Drug interactions and contraindications
- Clinical reasoning and differential diagnosis
Legal
- Civil vs criminal liability distinctions
- Constitutional concepts (habeas corpus, due process, equal protection)
- Contract law — offer, acceptance, consideration, breach, remedies
- Tort law fundamentals — negligence, strict liability
- Statutory interpretation principles
Finance
- Financial ratios (ROE, ROA, P/E, EV/EBITDA, current ratio)
- Fixed income — duration, convexity, yield curve dynamics
- Options strategies (calls, puts, collars, straddles, spreads)
- Accounting principles — GAAP vs IFRS, revenue recognition, depreciation
- Market microstructure, portfolio theory, risk management
Limitations and responsible use
- Medical: Not a substitute for clinical judgment. Always consult a licensed physician before making any medical decision.
- Legal: Not legal advice. Consult a qualified attorney licensed in your jurisdiction.
- Finance: Not personalized financial advice. Consult a registered investment advisor.
- Model may hallucinate facts — treat all outputs as a starting point, not ground truth.
- Performance may degrade on highly specialized sub-domains under-represented in training data.
- Evaluated on held-out splits from the same distribution — real-world performance may vary.
- Training data reflects publicly available datasets and may contain biases.
Repository contents
rajveer43/gemma-4-E4B-medical-legal-finance-qa/
├── adapter_config.json # LoRA adapter configuration
├── adapter_model.safetensors # LoRA weights (~100 MB)
├── tokenizer.json # Gemma 4 tokenizer
├── tokenizer_config.json
├── special_tokens_map.json
├── eval_results.json # Post-training evaluation results
└── model-unsloth.Q4_K_M.gguf # Quantized for llama.cpp / M4 Mac (~3.5 GB)Author
Rajveer — AI/ML Engineer and Open Source Contributor
  
Fine-tuned as part of domain-specific LLM adaptation research. Part of an ongoing series of open-source ML experiments — feedback and contributions welcome!
Citation
If you use this model in your research or project, please cite:
@misc{rajveer43-gemma4-domain-qa-2026,
author = {Rajveer},
title = {Gemma 4 E4B Fine-tuned on Medical, Legal and Finance QA},
year = {2026},
publisher = {HuggingFace},
howpublished = {\url{https://huggingface.co/rajveer43/gemma-4-E4B-medical-legal-finance-qa}},
}License
This model adapter inherits the Apache 2.0 license from the base Gemma 4 model. Free for commercial and research use with attribution.
