CoolFace
Modelpublic

emiliogirard/llama-3.1-8b-counseling-lora

sourceHugging Facellama3.1updated 5mo agoView on Hugging Face
0likes11downloads
Model Card

Llama-3.1-8B-Instruct — Mental Health Counseling (QLoRA)

A LoRA adapter fine-tuned from meta-llama/Llama-3.1-8B-Instruct on the Amod/mental_health_counseling_conversations dataset.

The model is trained to respond in a conversational, CBT-style therapeutic voice — reflective listening, 30-120 word responses, no markdown formatting (intended for voice/TTS output).


Training details

SettingValue
Base modelmeta-llama/Llama-3.1-8B-Instruct
QuantizationNF4 via bitsandbytes (4-bit, double quant)
Compute dtypebfloat16
LoRA rank16
LoRA alpha16
LoRA dropout0.05
Target modulesq_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Trainable params41.9M (0.52% of 8B total)
DatasetAmod/mentalhealthcounseling_conversations (3,441 train / 71 eval)
Epochs3
Effective batch size16 (per-device 1 × grad accum 16)
Learning rate2e-4 (cosine schedule, 3% warmup)
Weight decay0.01
Optimizerpagedadamw8bit
Max sequence length1024
Gradient checkpointingEnabled (non-reentrant)
Training time~3 hours

Hardware

  • —NVIDIA DGX Spark (Blackwell GB10, 128 GB unified memory, sm_121)
  • —PyTorch 2.11.0+cu130
  • —transformers, peft, trl, bitsandbytes (latest)

Training ran entirely on local hardware — zero cloud GPU cost, zero data uploaded to third-party fine-tuning services.

System prompt used during training

You are a compassionate, evidence-based CBT therapist. Respond naturally with reflective listening, keep responses conversational (30-120 words).

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-3.1-8B-Instruct",
    torch_dtype="bfloat16",
    device_map="auto",
)
model = PeftModel.from_pretrained(base, "emiliogirard/llama-3.1-8b-counseling-lora")
tokenizer = AutoTokenizer.from_pretrained("emiliogirard/llama-3.1-8b-counseling-lora")

messages = [
    {"role": "system", "content": "You are a compassionate, evidence-based CBT therapist. Respond naturally with reflective listening, keep responses conversational (30-120 words)."},
    {"role": "user", "content": "I'm feeling overwhelmed at work and can't sleep."},
]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)
outputs = model.generate(inputs, max_new_tokens=200, temperature=0.7)
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))

Intended use

Demonstration / portfolio project showing QLoRA fine-tuning of a 8B chat model on a conversational-therapy dataset, run fully on-premises on Blackwell hardware.

Limitations

  • —Not a clinical tool. This model is not validated for therapeutic or diagnostic use. It is not a substitute for a licensed mental health professional.
  • —Single epoch of training — production-grade counseling models require extensive safety alignment (e.g., DPO on crisis preference pairs, Llama-Guard 3 integration for input/output filtering), which is NOT part of this demo.
  • —Response quality should be evaluated before any real-world application.
  • —No formal eval benchmarks were run beyond training loss monitoring.

Production-grade adaptation

For production use, the following are recommended and available on request:

  • —DPO safety training on crisis preference pairs (suicide, self-harm, psychosis escalation)
  • —Llama-Guard 3 wrapping at both input and output (sub-200ms latency)
  • —Streaming FastAPI deployment (SSE / WebSocket) compatible with TTS
  • —NVFP4 quantization for Blackwell-native tensor-core inference
  • —Optional EAGLE-3 speculative decoding head for 3-4× throughput

Contact

Portfolio: pyloxforge.com Custom on-premises LLM fine-tuning and deployment available for healthcare, legal, and privacy-sensitive industries.