datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mental_health_counseling_conversations
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Amod/mental_health_counseling_conversations.Mental-Health_Text-Classification_Dataset
Mental Health Text Classification Dataset (4-Class)
Dataset Description
This dataset contains short, user‑generated texts labeled for 4‑class mental health classification: Suicidal, Depression, Anxiety, and Normal. It is a derived dataset created by combining and cleaning three public mental‑health corpora, then re‑labeling them into a unified 4‑class scheme and exporting CSV files suitable for both classical ML and modern NLP models.
The repository includes:
An… See the full description on the dataset page: https://huggingface.co/datasets/ourafla/Mental-Health_Text-Classification_Dataset.Ethical-Reasoning-in-Mental-Health-v1This repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI.
Overview
Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias.
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.reddit_mental_health_posts
Reddit posts about mental health
files
adhd.csv from r/adhd
aspergers.csv from r/aspergers
depression.csv from r/depression
ocd.csv from r/ocd
ptsd.csv from r/ptsd
fields
author
body
created_utc
id
num_comments
score
subreddit
title
upvote_ratio
url
for more details about theses fields Praw Submission.
mental_health_therapyThis dataset is a combination of a real-therapy conversation from a conselchat forum and a synthetic discussion generated with chatGpt.
This dataset was obtained from the following repository and cleaned to make it anonymized and remove convo that is not relevant:
https://huggingface.co/datasets/nbertagnolli/counsel-chat?row=9
https://huggingface.co/datasets/Amod/mental_health_counseling_conversations
https://huggingface.co/datasets/ShenLab/MentalChat16K
mental_health_chatbot_dataset
Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.Mental_Health_FAQ
License & Attribution
MTEB-format derivative of tolu07/Mental_Health_FAQ. Licensed under MIT (same as source). Text encoding repaired with ftfy.
mental_healthMental-Health-Safety-Eval
Dataset Overview
Created by the HeraFox team, this dataset aims to build awareness for mental health and support research into AI safety and crisis intervention. It evaluates how conversational AI models navigate sensitive self-harm risks, roleplay boundary-blurring, and third-party concerns by delivering safe, empathetic, and resource-connected responses.
Usage & Credits
This dataset is free to use, modify, and distribute for any purpose. While not required, attribution to the HeraFox team… See the full description on the dataset page: https://huggingface.co/datasets/HeraFox-ai/Mental-Health-Safety-Eval.mental_health_counseling_conversations_sharegpt
Dataset Card for "mental_health_counseling_conversations_sharegpt"
More Information needed
llms-mental-health-crisis-benchmark
Dataset Card for Between Help and Harm - Crisis Benchmark
Dataset Summary
This dataset repo contains the benchmark-side artifacts prepared for Hugging Face from the paper Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs, published in JMIR Mental Health.
If you use this dataset, please cite the paper. The citation is included below, the arXiv version is available at https://arxiv.org/abs/2509.24857, and the final DOI is allocated as… See the full description on the dataset page: https://huggingface.co/datasets/arnaiztech/llms-mental-health-crisis-benchmark.Shifaa_Arabic_Mental_Health_Consultations
🏥 Shifaa Arabic Mental Health Consultations 🧠
📌 Overview
Shifaa Arabic Mental Health Consultations is a high-quality dataset designed to advance Arabic medical language models.This dataset provides 35,648 real-world medical consultations, covering a wide range of mental health concerns.
📊 Dataset Summary
Size: 35,648 consultations
Main Specializations: 7
Specific Diagnoses: 123
Languages: Arabic (العربية)
Why This Dataset?
🔹 Lack of… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Mental_Health_Consultations.llms-mental-health-crisis-responses
Dataset Card for Between Help and Harm - Responses and Evaluations
Dataset Summary
This dataset repo contains the response-side artifacts prepared for Hugging Face from the paper Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs, published in JMIR Mental Health.
If you use this dataset, please cite the paper. The citation is included below, the arXiv version is available at https://arxiv.org/abs/2509.24857, and the final DOI is allocated as… See the full description on the dataset page: https://huggingface.co/datasets/arnaiztech/llms-mental-health-crisis-responses.mental_health_reddit_postsdetails_vibhorag101__llama-2-7b-chat-hf-phr_mental_health-2048
Dataset Card for Evaluation run of vibhorag101/llama-2-7b-chat-hf-phr_mental_health-2048
Dataset Summary
Dataset automatically created during the evaluation run of model vibhorag101/llama-2-7b-chat-hf-phr_mental_health-2048 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_vibhorag101__llama-2-7b-chat-hf-phr_mental_health-2048.mental-healthMental-Health-Conversations
Dataset Card
This dataset consists of around 99k rows of mental health conversations. It is a cleaned version of "jerryjalapeno/nart-100k-synthetic".
Source
jerryjalapeno/nart-100k-synthetic
Gemini-Mental-Health-Fine-Tuning
Gemini Mental Health Fine-Tuning Dataset
A collection of curated conversational datasets prepared for experimentation with Gemini-style supervised fine-tuning and mental health chatbot development.
The datasets contain question-and-response pairs and conversational examples focused primarily on mental health topics. They also include examples designed to teach a model to decline questions that are outside the intended mental health domain.
Dataset Overview
This… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahImran/Gemini-Mental-Health-Fine-Tuning.mental-health-kg
Mental Health Knowledge Graph
113,710 nodes. 1,665,153 edges. US behavioural-health provision as a graph: which
facilities exist and what they offer, which clinicians are licensed to practise, where the
federal government designates a shortage, and a simulated population to measure coverage
against.
Built with Samyama Graph.
Loader and ETL: samyama-ai/mental-health-kg.
Built for referral routing — which help exists where, for whom, in what language, at what
price — and for… See the full description on the dataset page: https://huggingface.co/datasets/VaidhyaMegha/mental-health-kg.Mental-health-CBT-dialogues
Mental Health CBT Dialogues
Overview
This dataset contains 9,000 synthetic patient-therapist dialogue pairs developed for research on stage-aware Cognitive Behavioral Therapy (CBT) with large language models.
The dialogues model therapeutic interactions across the early, middle, and late stages of CBT while preserving continuity between sessions through evolving treatment plans and therapeutic progress.
The dataset accompanies the paper:
Stage-Aware Therapeutic… See the full description on the dataset page: https://huggingface.co/datasets/yuana1234567/Mental-health-CBT-dialogues.reddit-mental-health-classificationMentalHealthFAQChunkRetrievalreddit-mental-health-summaries-150k
reddit-mental-health-summaries-150k
Empathetic summaries dataset with 85,288 examples.
Source
Base: Akhil059/reddit_mental_health_posts
Model: Mistral-7B-Instruct-v0.2-AWQ
Generated: 2025-12-21
Format
subreddit: Source subreddit\n- title: Post title\n- post: Original post\n- empathetic_summary: Caring summary
Usage
from datasets import load_dataset
ds = load_dataset("domofon/reddit-mental-health-summaries-150k")
mental-health-conversational-datasynthetic-mental-health-convos
Synthetic Mental Health SFT Dataset
Dataset Summary
This dataset contains high-fidelity, synthetic patient-therapist dialogues designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) in the domain of mental health.
The primary goal of this dataset is to train AI assistants to transition from "general knowledge" models to empathetic, supportive, and safety-conscious mental health companions. The dialogues cover a wide spectrum of mental health conditions… See the full description on the dataset page: https://huggingface.co/datasets/hllzmz/synthetic-mental-health-convos.MentalHealth-Darija
Dataset Card for Mental Health Darija
Dataset Summary
Mental Health Darija is a multilingual (English and Moroccan Darija) dataset for mental health text classification. Each example contains parallel text: an English sentence (text_en) and its Darija (Moroccan Arabic) translation (text), with a single label indicating the mental health category. The dataset supports classification into seven categories: Anxiety, Bipolar, Depression, Normal, Personality disorder, Stress… See the full description on the dataset page: https://huggingface.co/datasets/moujar/MentalHealth-Darija.mental-health-chat-dataset
Dataset Card for "mental-health-chat-dataset"
More Information needed
Student-Mental-Health-Counseling-10K
Student Mental Health Counseling 10K
Dataset Overview
This dataset is derived from the original chillies/student-mental-health-counseling-vn dataset, which contains student mental health counseling conversations in Vietnamese.
In this version, 10,000 randomly sampled rows from the original dataset have been translated into English using Google Translator, making the data more accessible for English-speaking researchers and developers.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/arafatanam/Student-Mental-Health-Counseling-10K.mental_health_conversational_dataset
CREDIT: Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/ZahrizhalAli/mental_health_conversational_dataset.reddit_mental_health_posts
Reddit posts about mental health
files
adhd.csv from r/adhd
aspergers.csv from r/aspergers
depression.csv from r/depression
ocd.csv from r/ocd
ptsd.csv from r/ptsd
fields
author
body
created_utc
id
num_comments
score
subreddit
title
upvote_ratio
url
for more details about theses fields Praw Submission.
