CoolFace
Modelpublic

rajveer43/gemma-4-E4B-medical-legal-finance-qa

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
13likes260downloads
Model Card

Gemma 4 E4B — Medical · Legal · Finance Domain QA

Fine-tuned version of `google/gemma-4-E4B-it` across three professional domains — Medical, Legal, and Finance — using QLoRA (4-bit NF4) with Optuna-tuned hyperparameters, trained on Kaggle T4 GPU.

Trained by rajveer43  ·  Apache 2.0 license


What this model does

This model answers domain-specific professional questions across three verticals. Domain context is passed via the system prompt — the same adapter handles all three.

DomainTraining datasetExample queries
Medicalmedalpaca/medical_meadow_medqaDrug mechanisms, diagnostics, clinical reasoning
Legalnguyen-brat/legal_case_qaCase law, contract interpretation, legal concepts
Financegbharti/finance-alpacaMarket instruments, valuation, accounting concepts

Model details

PropertyValue
Base modelgoogle/gemma-4-E4B-it
Fine-tuning methodQLoRA — BitsAndBytes 4-bit NF4 + LoRA adapters
LoRA rank / alphar=16 · alpha=32 · RSLoRA
Target modulesq\proj, k\proj, v\proj, o\proj, gate\proj, up\proj, down\_proj
Trainable parameters~1% of total (base frozen in 4-bit)
Training hardwareKaggle T4 GPU (15 GB VRAM)

Training pipeline

medical_meadow_medqa + legal_case_qa + finance-alpaca
               │
               ▼
    Format to Gemma 4 chat template
    (system prompt encodes domain context)
               │
               ▼
    google/gemma-4-E4B-it
    + BitsAndBytes 4-bit NF4
    + LoRA adapters (r=16, RSLoRA, ~1% params)
               │
               ▼
    Optuna HPO — 8 trials × 100 steps
    TPE sampler + MedianPruner
    Search: lr, batch, grad_accum, warmup
               │
               ▼
    SFTTrainer — 3 epochs, cosine LR
    best checkpoint saved per epoch
               │
               ▼
    GGUF Q4_K_M → llama.cpp

Stack: unsloth · trl · peft · bitsandbytes · accelerate · optuna · datasets


Quickstart

Option A — Unsloth (fastest)

python
from unsloth import FastModel
import torch

model, tokenizer = FastModel.from_pretrained(
    model_name   = "rajveer43/gemma-4-E4B-medical-legal-finance-qa",
    load_in_4bit = True,
)
FastModel.for_inference(model)

SYSTEM_PROMPTS = {
    "medical": (
        "You are a highly knowledgeable medical AI assistant. "
        "Answer clinical and biomedical questions accurately. "
        "Always advise consulting a licensed physician for personal medical decisions."
    ),
    "legal": (
        "You are an expert legal AI assistant. "
        "Provide accurate information about laws, precedents, and legal concepts. "
        "Always clarify responses are not a substitute for licensed legal counsel."
    ),
    "finance": (
        "You are a knowledgeable financial AI assistant. "
        "Answer questions about markets, instruments, and economic concepts accurately. "
        "Remind users this is not personalized financial advice."
    ),
}

def ask(question: str, domain: str = "medical", max_new_tokens: int = 256) -> str:
    messages = [
        {"role": "system", "content": SYSTEM_PROMPTS[domain]},
        {"role": "user",   "content": question},
    ]

    prompt = tokenizer.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True,
    )

    # FIX: Bypass the Unsloth-patched multimodal processor entirely.
    # For text-only inference, call the underlying fast tokenizer directly
    # so the image processor pipeline is never triggered.
    underlying_tokenizer = tokenizer.tokenizer  # unwrap: Gemma4Processor -> PreTrainedTokenizerFast

    inputs = underlying_tokenizer(
        prompt,
        return_tensors="pt",
        padding=True,
    ).to(model.device)

    with torch.no_grad():
        output = model.generate(
            **inputs,
            max_new_tokens=max_new_tokens,
            temperature=0.1,
            top_p=0.9,
            do_sample=True,
        )

    input_len = inputs["input_ids"].shape[-1]
    return underlying_tokenizer.decode(output[0][input_len:], skip_special_tokens=True)

# Medical
print(ask("What is the first-line treatment for Type 2 diabetes?", domain="medical"))
# Legal
print(ask("What constitutes a breach of contract under common law?", domain="legal"))
# Finance
print(ask("What is the difference between ROE and ROA?", domain="finance"))

Option B — PEFT + Transformers

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
import torch

bnb_config = BitsAndBytesConfig(
    load_in_4bit              = True,
    bnb_4bit_quant_type       = "nf4",
    bnb_4bit_compute_dtype    = torch.bfloat16,
    bnb_4bit_use_double_quant = True,
)
base = AutoModelForCausalLM.from_pretrained(
    "google/gemma-4-E4B-it",
    quantization_config = bnb_config,
    device_map          = "auto",
)
model     = PeftModel.from_pretrained(base, "rajveer43/gemma-4-E4B-medical-legal-finance-qa")
tokenizer = AutoTokenizer.from_pretrained("rajveer43/gemma-4-E4B-medical-legal-finance-qa")

Option C — GGUF on M4 Mac (llama.cpp + Metal)

bash
# Install
brew install llama.cpp

# Download GGUF
huggingface-cli download \
  rajveer43/gemma-4-E4B-medical-legal-finance-qa \
  --include "*.gguf" \
  --local-dir ~/models/gemma4-domain-qa

# Interactive chat
llama-cli \
  -m ~/models/gemma4-domain-qa/model-unsloth.Q4_K_M.gguf \
  -ngl 99 \
  --chat-template gemma \
  -p "You are a knowledgeable medical AI assistant." \
  -i

# OpenAI-compatible server
llama-server \
  -m ~/models/gemma4-domain-qa/model-unsloth.Q4_K_M.gguf \
  -ngl 99 \
  --host 0.0.0.0 --port 8080

Expected on M4 Mac: 35–50 tokens/sec · ~4–5 GB unified memory


Domain coverage

Medical

  • —First-line treatment protocols and clinical guidelines
  • —Drug mechanism of action and pharmacology
  • —Diagnostic criteria (sepsis, diabetes, hypertension, etc.)
  • —Drug interactions and contraindications
  • —Clinical reasoning and differential diagnosis

Legal

  • —Civil vs criminal liability distinctions
  • —Constitutional concepts (habeas corpus, due process, equal protection)
  • —Contract law — offer, acceptance, consideration, breach, remedies
  • —Tort law fundamentals — negligence, strict liability
  • —Statutory interpretation principles

Finance

  • —Financial ratios (ROE, ROA, P/E, EV/EBITDA, current ratio)
  • —Fixed income — duration, convexity, yield curve dynamics
  • —Options strategies (calls, puts, collars, straddles, spreads)
  • —Accounting principles — GAAP vs IFRS, revenue recognition, depreciation
  • —Market microstructure, portfolio theory, risk management

Limitations and responsible use

  • —Medical: Not a substitute for clinical judgment. Always consult a licensed physician before making any medical decision.
  • —Legal: Not legal advice. Consult a qualified attorney licensed in your jurisdiction.
  • —Finance: Not personalized financial advice. Consult a registered investment advisor.
  • —Model may hallucinate facts — treat all outputs as a starting point, not ground truth.
  • —Performance may degrade on highly specialized sub-domains under-represented in training data.
  • —Evaluated on held-out splits from the same distribution — real-world performance may vary.
  • —Training data reflects publicly available datasets and may contain biases.

Repository contents

rajveer43/gemma-4-E4B-medical-legal-finance-qa/
├── adapter_config.json           # LoRA adapter configuration
├── adapter_model.safetensors     # LoRA weights (~100 MB)
├── tokenizer.json                # Gemma 4 tokenizer
├── tokenizer_config.json
├── special_tokens_map.json
├── eval_results.json             # Post-training evaluation results
└── model-unsloth.Q4_K_M.gguf    # Quantized for llama.cpp / M4 Mac (~3.5 GB)

Author

Rajveer — AI/ML Engineer and Open Source Contributor

![GitHub](https://github.com/rajveer43) ![HuggingFace](https://huggingface.co/rajveer43) ![LinkedIn](https://linkedin.com/in/rajveer43)

Fine-tuned as part of domain-specific LLM adaptation research. Part of an ongoing series of open-source ML experiments — feedback and contributions welcome!


Citation

If you use this model in your research or project, please cite:

bibtex
@misc{rajveer43-gemma4-domain-qa-2026,
  author       = {Rajveer},
  title        = {Gemma 4 E4B Fine-tuned on Medical, Legal and Finance QA},
  year         = {2026},
  publisher    = {HuggingFace},
  howpublished = {\url{https://huggingface.co/rajveer43/gemma-4-E4B-medical-legal-finance-qa}},
}

License

This model adapter inherits the Apache 2.0 license from the base Gemma 4 model. Free for commercial and research use with attribution.