datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mental_health_counseling_conversations
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Amod/mental_health_counseling_conversations.Ethical-Reasoning-in-Mental-Health-v1This repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI.
Overview
Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias.
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.mental_health_chatbot_dataset
Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.phr_mental_therapy_dataset
Dataset Card for "phr_mental_health_dataset"
This dataset is a cleaned version of nart-100k-synthetic
The data is generated synthetically using gpt3.5-turbo using this script.
The dataset had a "sharegpt" style JSONL format, with each JSON having keys "human" and "gpt", having an equal number of both.
The data was then cleaned, and the following changes were made
The names "Alex" and "Charlie" were removed from the dataset, which can often come up in the conversation of fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/vibhorag101/phr_mental_therapy_dataset.Mental-Health-Conversations
Dataset Card
This dataset consists of around 99k rows of mental health conversations. It is a cleaned version of "jerryjalapeno/nart-100k-synthetic".
Source
jerryjalapeno/nart-100k-synthetic
reddit-mental-health-summaries-150k
reddit-mental-health-summaries-150k
Empathetic summaries dataset with 85,288 examples.
Source
Base: Akhil059/reddit_mental_health_posts
Model: Mistral-7B-Instruct-v0.2-AWQ
Generated: 2025-12-21
Format
subreddit: Source subreddit\n- title: Post title\n- post: Original post\n- empathetic_summary: Caring summary
Usage
from datasets import load_dataset
ds = load_dataset("domofon/reddit-mental-health-summaries-150k")
Mental-health-CBT-dialogues
Mental Health CBT Dialogues
Overview
This dataset contains 9,000 synthetic patient-therapist dialogue pairs developed for research on stage-aware Cognitive Behavioral Therapy (CBT) with large language models.
The dialogues model therapeutic interactions across the early, middle, and late stages of CBT while preserving continuity between sessions through evolving treatment plans and therapeutic progress.
The dataset accompanies the paper:
Stage-Aware Therapeutic… See the full description on the dataset page: https://huggingface.co/datasets/yuana1234567/Mental-health-CBT-dialogues.Gemini-Mental-Health-Fine-Tuning
Gemini Mental Health Fine-Tuning Dataset
A collection of curated conversational datasets prepared for experimentation with Gemini-style supervised fine-tuning and mental health chatbot development.
The datasets contain question-and-response pairs and conversational examples focused primarily on mental health topics. They also include examples designed to teach a model to decline questions that are outside the intended mental health domain.
Dataset Overview
This… See the full description on the dataset page: https://huggingface.co/datasets/AbdullahImran/Gemini-Mental-Health-Fine-Tuning.synthetic-mental-health-convos
Synthetic Mental Health SFT Dataset
Dataset Summary
This dataset contains high-fidelity, synthetic patient-therapist dialogues designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) in the domain of mental health.
The primary goal of this dataset is to train AI assistants to transition from "general knowledge" models to empathetic, supportive, and safety-conscious mental health companions. The dialogues cover a wide spectrum of mental health conditions… See the full description on the dataset page: https://huggingface.co/datasets/hllzmz/synthetic-mental-health-convos.python-mental-execution-traces
Python Mental Execution Traces
A 12,000-row prompt/completion dataset for evaluating and training language models to mentally execute self-contained Python 3 snippets without running them. Completions provide the expected standard output together with a concise variable trace or explanation.
Dataset structure
The JSONL file contains two text fields:
prompt: a Python mental-execution problem.
completion: the expected stdout and concise reasoning or variable trace.… See the full description on the dataset page: https://huggingface.co/datasets/ILoveBuns/python-mental-execution-traces.Mental_Health_FAQContent
Mental health includes our emotional, psychological, and social well-being. Mental health is integral to living a healthy, balanced life. It affects how we think, feel, and act. It also helps determine how we handle stress, relate to others, and make choices. Emotional and mental health is important because it’s a vital part of your life and impacts your thoughts, behaviors and emotions. Being healthy emotionally can promote productivity and effectiveness in activities like work… See the full description on the dataset page: https://huggingface.co/datasets/tolu07/Mental_Health_FAQ.mental_health_conversational_dataset
CREDIT: Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/ZahrizhalAli/mental_health_conversational_dataset.mental-model-of-non-linear-promptingDownload Research Paper PDF
Mental model of Non-linear prompting:
Self-organization vs. Self-assembly
nature of AI
Tomaž Flegar
Institute for applied consciousness research
March the 30th, 2026
tomazf8@gmail.com
Primary Keywords: Self-organization, self-assembly, non-linear dynamics,
hyper-dimensional matrix, transformer architecture, coherentive communication,
phenomenological language, emergent complexity, attractor geometry, crystallization.
Secondary… See the full description on the dataset page: https://huggingface.co/datasets/tomazf8/mental-model-of-non-linear-prompting.mental_health_counseling_conversations
Amod/mental_health_counseling_conversations
This data is cloned from https://huggingface.co/datasets/Amod/mental_health_counseling_conversations
Dataset Summary
This dataset is a collection of questions and answers sourced from two online counseling and therapy platforms. The questions cover a wide range of mental health topics, and the answers are provided by qualified psychologists. The dataset is intended to be used for fine-tuning language models to improve their… See the full description on the dataset page: https://huggingface.co/datasets/MaggiePai/mental_health_counseling_conversations.mental-spaces
Mental Spaces Corpus
Version: 0.1.0
The Mental Spaces Corpus is a controlled suite of natural-language stimuli for testing
whether language models keep base-space and alternative-space discourse targets
separate. It is designed for probing, causal interventions, and behavioral readouts in
mental-space constructions such as counterfactuals, belief contexts, and depictive
spaces, including nested belief and nested depictive spaces.
This release is a stimulus suite for controlled… See the full description on the dataset page: https://huggingface.co/datasets/osteele/mental-spaces.Student-Mental-Health-Counseling-EN
Student Mental Health Counseling (EN)
Dataset Overview
This dataset is the final cleaned and filtered version in a three-step pipeline
that started with the original Vietnamese counseling dataset.
Step
Dataset
Rows
1. Original
chillies/student-mental-health-counseling-vn
750,169
2. Translated
arafatanam/Student-Mental-Health-Counseling-750K
750,169
3. Filtered (this dataset)
arafatanam/Student-Mental-Health-Counseling-EN
52,254
The goal of this… See the full description on the dataset page: https://huggingface.co/datasets/arafatanam/Student-Mental-Health-Counseling-EN.MentalHealth-Support
Important Note
This dataset is created from merging two datasets from different sources and has been formatted according to the "messages", "role", "content" chat format. I do not claim any ownership of this dataset.
Keep in mind that this dataset is entirely synthetic. It is not fully representative of real therapy situations. If you are training an LLM therapist keep in mind the limitations of LLMs and highlight those limitations to users in a responsible manner.
Since Mental… See the full description on the dataset page: https://huggingface.co/datasets/ShivomH/MentalHealth-Support.mental_health_counseling_conversations_koAmod/mental_health_counseling_conversations
GPT-4o 를 사용해 한글로 번역한 데이터셋입니다.
mental_health_datasetmental_health_psychology_curated_alpacaThis dataset is primarily composed of multiple choice questions on mental health, psychological issues, and training thereof. The questions are geared toward a professional training for a certification.
mental-health-ptMental-Health-Couseling
Mental Health Counseling Conversations (Cleaned)
Dataset Overview
This dataset is derived from the original Amod/mental_health_counseling_conversations dataset, which contains mental health counseling conversations.
In this version, duplicate Context-Response pairs have been removed to improve data quality and usability.
Dataset Details
Dataset Name: arafatanam/Mental-Health-Counseling
Source Dataset: Amod/mental_health_counseling_conversations
Modifications:… See the full description on the dataset page: https://huggingface.co/datasets/arafatanam/Mental-Health-Couseling.mental-health-conversation-sft-12k
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/DenCT/mental-health-conversation-sft-12k.Medical-and-Mental-Health
Dataset Card
These datasets were obtained from the sources mentioned below. These datasets were customized and modified to share the same file and data format specifically for the purpose of fine-tuning LLMs.
As these datasets contain general medical and mental health data, it is exptected that the datasets will be used responsibly.
Sources
FunPang/medical_dataset
jerryjalapeno/nart-100k-synthetic
fadodr/mental_health_therapy
marmikpandya/mental-health
mental_health_counseling_conversations
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/alexjoseph0905/mental_health_counseling_conversations.mental_health_counseling_responses
Dataset Card for Mental Health Counseling Responses
This dataset contains responses to questions from mental health counseling sessions.
The responses are rated by LLMs using the dimensions: empathy, appropriateness, and relevance.
A detailed explanation of the rating process can be found in this blog post.
For a detailed analysis of LLM-generated responses and their comparison to human responses, refer to this blog post.
The original data with the human responses can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_responses.pashto-mental-health-counseling-3k
🧠 Pashto Mental Health Counseling 3K
This dataset is a specialized collection of 3,000 conversational pairs focused on mental health counseling, translated and culturally adapted into Pashto. It is designed to train LLMs to provide empathetic, supportive, and culturally relevant responses in a therapeutic context.
🌟 Overview
Mental health resources in Pashto are scarce. This dataset aims to bridge that gap by providing high-quality counseling dialogues. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-mental-health-counseling-3k.mental_health_Chatbot
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Iamzoo/mental_health_Chatbot.mental_health_counseling_conversations_rated
Dataset Card for Mental Health Counseling Conversations Rated
This dataset extends the existing dataset Mental Health Counseling Conversations and adds ratings for the responses.
Dataset Details
This dataset is an extension for the dataset Mental Health Counseling Conversations.
It adds ratings for the responses generated by four different LLMs. The responses are rated across the following dimensions:
empathy
appropriateness
relevance
The following four LLMs are used… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_conversations_rated.Mental_Health_Support_ChatBOT_Conversation
Mental Health Support Dataset
Instruction–response pairs for training supportive, non-diagnostic,
safety-aware mental health chatbots.
Fields
instruction: user message
response: Bot reposne
category: intent label
Safety
This dataset includes crisis escalation examples and refusal patterns.
Not a replacement for professional care.
