emiliogirard/llama-3.1-8b-counseling-lora
Llama-3.1-8B-Instruct — Mental Health Counseling (QLoRA)
A LoRA adapter fine-tuned from meta-llama/Llama-3.1-8B-Instruct on the Amod/mental_health_counseling_conversations dataset.
The model is trained to respond in a conversational, CBT-style therapeutic voice — reflective listening, 30-120 word responses, no markdown formatting (intended for voice/TTS output).
Training details
Hardware
- NVIDIA DGX Spark (Blackwell GB10, 128 GB unified memory, sm_121)
- PyTorch 2.11.0+cu130
- transformers, peft, trl, bitsandbytes (latest)
Training ran entirely on local hardware — zero cloud GPU cost, zero data uploaded to third-party fine-tuning services.
System prompt used during training
You are a compassionate, evidence-based CBT therapist. Respond naturally with reflective listening, keep responses conversational (30-120 words).
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B-Instruct",
torch_dtype="bfloat16",
device_map="auto",
)
model = PeftModel.from_pretrained(base, "emiliogirard/llama-3.1-8b-counseling-lora")
tokenizer = AutoTokenizer.from_pretrained("emiliogirard/llama-3.1-8b-counseling-lora")
messages = [
{"role": "system", "content": "You are a compassionate, evidence-based CBT therapist. Respond naturally with reflective listening, keep responses conversational (30-120 words)."},
{"role": "user", "content": "I'm feeling overwhelmed at work and can't sleep."},
]
inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)
outputs = model.generate(inputs, max_new_tokens=200, temperature=0.7)
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))Intended use
Demonstration / portfolio project showing QLoRA fine-tuning of a 8B chat model on a conversational-therapy dataset, run fully on-premises on Blackwell hardware.
Limitations
- Not a clinical tool. This model is not validated for therapeutic or diagnostic use. It is not a substitute for a licensed mental health professional.
- Single epoch of training — production-grade counseling models require extensive safety alignment (e.g., DPO on crisis preference pairs, Llama-Guard 3 integration for input/output filtering), which is NOT part of this demo.
- Response quality should be evaluated before any real-world application.
- No formal eval benchmarks were run beyond training loss monitoring.
Production-grade adaptation
For production use, the following are recommended and available on request:
- DPO safety training on crisis preference pairs (suicide, self-harm, psychosis escalation)
- Llama-Guard 3 wrapping at both input and output (sub-200ms latency)
- Streaming FastAPI deployment (SSE / WebSocket) compatible with TTS
- NVFP4 quantization for Blackwell-native tensor-core inference
- Optional EAGLE-3 speculative decoding head for 3-4× throughput
Contact
Portfolio: pyloxforge.com Custom on-premises LLM fine-tuning and deployment available for healthcare, legal, and privacy-sensitive industries.
