CoolFace
Modelpublic

andong90/DeepSeek-R1-Distill-Qwen-7B-student-mental-health-json

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes12downloads
Model Card

Model Card for Model ID

<!-- Provide a quick summary of what the model is/does. -->

Model Details

Model Description

<!-- Provide a longer summary of what this model is. -->

This is the model card of a ๐Ÿค— transformers model that has been pushed on the Hub. This model card has been automatically generated.

  • โ€”Developed by: [an.dong90]
  • โ€”Model type: [Fine tuned distilled Deepseek R1 Qwen 7B model]
  • โ€”Language(s) (NLP): [English]
  • โ€”License: [MIT]
  • โ€”**Finetuned from model [deepseek-ai/DeepSeek-R1-Distill-Qwen-7B]

Model Sources [optional]

<!-- Provide the basic links for the model. -->

Traning process: https://github.com/dojian/mentalhealthchatbot/blob/main/notebooks/Json_conv.ipynb

Uses

<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->

prompttest = """Given a student's Conversation History and Current Message, extract the relevant metadata, including emotion type, emotion intensity (1-5), problem type, and counseling strategy. Then answer the student's Current Message as a counselor based on the metadata. Keep it concise but affirmative. The counselor must return a Structured JSON Response with these fields: "emotiontype","emotionintensity", "problemtype", "counseling_strategy","answer".

Student:

Conversation History: {user_history}

Current Message: {user_text}

Counselor Structured JSON Response:

"""

FastLanguageModel.forinference(model) inputs = tokenizer([prompttest.format(userhistory=userhistory,usertext=usertext)], returntensors="pt").to("cuda") outputs = model.generate( inputids=inputs.inputids, attentionmask=inputs.attentionmask, maxnewtokens=250, eostokenid=tokenizer.eostokenid, numreturnsequences=1, temperature=0.6, # deepseek doc recommended 0.6 to balance creativity and coherence, avoiding repetitive or nonsensical outputs. topp=0.9, # Reduces repeated phrases usecache=True, ) response = tokenizer.decode(outputs[0],skipspecial_tokens=True)

[More Information Needed]

Training Details

Training Data

<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. --> https://huggingface.co/datasets/jordanfan/esconv_processed

Training Procedure

<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->

Preprocessing [optional]

[More Information Needed]

Training Hyperparameters
  • โ€”Training regime: [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
Speeds, Sizes, Times [optional]

<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->

[More Information Needed]

Evaluation

<!-- This section describes the evaluation protocols and provides the results. --> 3 LLMs as judges DeepSeek R1 Distilled Llama 8B DeepSeek R1 Distilled Qwen 7B Mistral 7B v0.3

Assessed generated responses based on empathy, appropriateness, and relevance on scale of 1-5 Metrics proposed on Medium article in similar mental health setting*

Averaged score across judges

Median Empathy-4.00 Appropriateness-5.00 Relevance-4.33

Testing Data, Factors & Metrics

Testing Data

<!-- This should link to a Dataset Card if possible. -->

[More Information Needed]

Factors

<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->

[More Information Needed]

Metrics

<!-- These are the evaluation metrics being used, ideally with a description of why. -->

[More Information Needed]

Results

[More Information Needed]

Summary