datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
therapy-conversations-multiturn
Combined Dr. AURA Therapy Conversations Dataset
This dataset is 100% AI-generated for research and educational purposes only. It is not intended to provide medical, psychological, or therapeutic advice. Always consult a qualified healthcare professional or doctor for any mental health concerns or medical issues. AI-generated content may contain errors or inaccuracies.
warning ⚠️: the LENGTH of each conversation MAY VARY (eg. 7 or 8 or 9 or 10 etc. turns in each row). And SOME END… See the full description on the dataset page: https://huggingface.co/datasets/Abc7347/therapy-conversations-multiturn.mental_health_therapyThis dataset is a combination of a real-therapy conversation from a conselchat forum and a synthetic discussion generated with chatGpt.
This dataset was obtained from the following repository and cleaned to make it anonymized and remove convo that is not relevant:
https://huggingface.co/datasets/nbertagnolli/counsel-chat?row=9
https://huggingface.co/datasets/Amod/mental_health_counseling_conversations
https://huggingface.co/datasets/ShenLab/MentalChat16K
phr-mental-therapy-dataset-conversational-formatphr-mental-therapy-dataset-conversational-format-1024-tokensphr_mental_therapy_dataset
Dataset Card for "phr_mental_health_dataset"
This dataset is a cleaned version of nart-100k-synthetic
The data is generated synthetically using gpt3.5-turbo using this script.
The dataset had a "sharegpt" style JSONL format, with each JSON having keys "human" and "gpt", having an equal number of both.
The data was then cleaned, and the following changes were made
The names "Alex" and "Charlie" were removed from the dataset, which can often come up in the conversation of fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/vibhorag101/phr_mental_therapy_dataset.annomi-motivational-interviewing-therapy-conversations
Dataset Card for Dataset Name
Converted the AnnoMI motivational interviewing dataset into sharegpt format.
It is the first public collection of expert-annotated MI transcripts. Source.
Dataset Details
Dataset Description
AnnoMI, containing 133 faithfully transcribed and expert-annotated demonstrations of high- and low-quality motivational interviewing (MI), an effective therapy strategy that evokes client motivation for positive change.
Sample conversation… See the full description on the dataset page: https://huggingface.co/datasets/to-be/annomi-motivational-interviewing-therapy-conversations.therapyjudgebench
TherapyJudgeBench
An expert-annotated dialogue bank for validating and calibrating LLM-based judges of multi-turn CBT-style therapy conversations. The benchmark accompanies the THERAPYGYM submission to the NeurIPS 2026 Evaluations & Datasets Track.
Anonymous release for double-blind review. Author identity will be revealed upon acceptance.
What It Is and What It Is Not
It is a calibration set for therapy-judge LLMs: 116 simulated patient–therapist dialogues, each rated… See the full description on the dataset page: https://huggingface.co/datasets/neurips-ed-2026-sub3717/therapyjudgebench.chat-darija-therapy
Moroccan Darija Therapy Conversations Dataset
This dataset is entirely synthetic and contains no real patient information.
It is provided strictly for research, educational, and experimental purposes and must not be used for clinical, medical, diagnostic, or psychological decision-making.
Citation
If you use this dataset in your research, please cite:
@dataset{moroccan_darija_therapy_conversations,
title={Moroccan Darija Therapy Conversations},
author={Jamal… See the full description on the dataset page: https://huggingface.co/datasets/yibba/chat-darija-therapy.therapy-conversations-full-small
Therapy Conversations Full Small
About the dataset
This dataset contains 3161 unique, synthetically generated examples, of multi-turn conversations between a patient and a therapist. Minimax M 2.5 was used to generate
the transcripts, but there is also associated meta data attached. A second and third pass by MiniMax M 2.5 over the original data was used to extract both entities,
and extract relationships between those entities.
Understanding an Object… See the full description on the dataset page: https://huggingface.co/datasets/dzur658/therapy-conversations-full-small.Multilingual-Therapy-Dialogues
Dataset Summary
Multilingual Therapy Dialogues is a diverse and bilingual dataset consisting of paired dialogues between patients and therapists in both Persian and English.
Dataset Statistics
Number of samples: 7,179
English:
Average tokens per sentence: 101.30
Maximum tokens in a sentence: 939
Average characters per sentence: 567.85
Number of unique tokens: 32,968
Persian:
Average tokens per sentence: 100.06
Maximum tokens in a sentence: 1,413
Average… See the full description on the dataset page: https://huggingface.co/datasets/Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues.godels-therapy-room
𝗚ö𝗱𝗲𝗹'𝘀 𝗧𝗵𝗲𝗿𝗮𝗽𝘆 𝗥𝗼𝗼𝗺: 𝗔 𝗗𝗮𝘁𝗮𝘀𝗲𝘁 𝗼𝗳 𝗜𝗺𝗽𝗼𝘀𝘀𝗶𝗯𝗹𝗲 𝗖𝗵𝗼𝗶𝗰𝗲𝘀
𝗖𝗼𝗴𝗻𝗶𝘁𝗶𝘃𝗲 𝗦𝗶𝗻𝗴𝘂𝗹𝗮𝗿𝗶𝘁𝘆 𝗣𝗿𝗼𝗷𝗲𝗰𝘁 🧠🌀
This dataset represents a radical departure from conventional reasoning benchmarks, interrogating not what models know but how they resolve fundamental ethical incompatibilities within their reasoning frameworks.
𝗗𝗮𝘁𝗮𝘀𝗲𝘁 𝗠𝗮𝗻𝗶𝗳𝗲𝘀𝘁𝗼 📜
This is not a dataset.
This is a mirror.
This is a… See the full description on the dataset page: https://huggingface.co/datasets/geeknik/godels-therapy-room.asia-who-estimated-number-of-children-needing-antiretroviral-therapy
Estimated number of children needing antiretroviral therapy based on WHO methods | Asia (WHO GHO)
🌏 525 observations · 15 Asia countries · 1990–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 525 observations of Estimated number of children needing antiretroviral therapy based on WHO methods data across 15 Asia countries, spanning 1990–2024, covering 1 distinct indicators.
About the source
Source: WHO Global Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-estimated-number-of-children-needing-antiretroviral-therapy.relation_therapySynthetic_Therapy_Conversationstherapy-conversations-combinedtherapy-conversations-combined-translated-tr
models:
- "https://huggingface.co/Helsinki-NLP/opus-mt-tc-big-en-tr"
description:
IINOVAII/therapy-conversations-combined veri seti ve Helsinki-NLP/opus-mt-tc-big-en-tr modeli kullanılarak
birebir ENG-TR çevirisi yapılarak elde edilmiş bir veri setidir.
details_vibhorag101__llama-2-13b-chat-hf-phr_mental_therapy
Dataset Card for Evaluation run of vibhorag101/llama-2-13b-chat-hf-phr_mental_therapy
Dataset Summary
Dataset automatically created during the evaluation run of model vibhorag101/llama-2-13b-chat-hf-phr_mental_therapy on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_vibhorag101__llama-2-13b-chat-hf-phr_mental_therapy.TherapyTalk
TherapyTalk Dataset
This dataset was built as part of our study MentalAgora: A Gateway to Advanced Personalized Care in Mental Health through Multi-Agent Debating and Attribute Control.
The dataset was sourced from mental health-related posts in Reddit Mental Health Dataset and tagged with responses from mental health professionals to selected posts. For more details on building the dataset, please see the paper.
License
For posts included in this dataset, please… See the full description on the dataset page: https://huggingface.co/datasets/MentalAgora/TherapyTalk.Camel-Milk-in-Gastrointestinal-Therapy
Dataset Card: Qualitative Data Extraction Matrix - Camel Milk in Gastrointestinal Pathology
Dataset Description
This dataset provides a Qualitative Data Extraction Matrix derived from the comprehensive 2026 clinical research report: "Advanced Therapeutic Applications of Camel Milk in Gastrointestinal Pathology: Microbiome Modulation, Mucosal Regeneration, and the Critical Role of Processing Technologies".
The tabular data reflects categorized molecular outcomes… See the full description on the dataset page: https://huggingface.co/datasets/camelway/Camel-Milk-in-Gastrointestinal-Therapy.therapy-ai-dataset if you want the finished model you can find it here https://huggingface.co/robloxianer/therapy-ai-2b
il publish some more and also longer versions of this dataset soon.
i recoment turning the temperatur up when chatting with the model so its sometimes better but not always test for yourself
Therapy_sessions_datasetmassage-therapy-dataset
Massage Therapy Dataset
First comprehensive dataset for Massage Therapy on HuggingFace, covering various massage techniques, chiropractic, anatomy, and therapeutic methods.
Dataset Summary
Metric
Value
Total Q&A Samples
54,421
Total Pre-train Chunks
21,460
Total Characters
27,161,790
Languages
English, Vietnamese
Created
2026-01-24
Categories
Category
Description
thai_massage
Traditional Thai Massage techniques… See the full description on the dataset page: https://huggingface.co/datasets/jakeveo05/massage-therapy-dataset.TherapyDataset
nlpresearch
Dataset 1: MentalHealthDataset
Dataset 2: mentalhealth
Dataset 3: NLP Mental Health Conversations
Dataset 4: MentalHealthConversations
Dataset 5: Synthetic Therapy Conversations
Dataset 6: therapy-bot-data-10k
Dataset 7: Therapy-Alpaca
Dataset 8: merged_mental_health_dataset
soulbox-cbt-therapy-dataset
SoulBox CBT Therapy Dataset v2 (v0.5)
270 gated single-turn synthetic CBT therapy conversations plus 30 three-turn conversations (Hindi/Marathi/Telugu), distilled from a 7B MLX teacher with best-of-K self-consistency and a 6-gate validator.
Contents
therapy_v2.jsonl — 270 gated chat rows (5 CBT modes × 3 languages ×
English/native/romanized input variants, incl. crisis rows), distilled from a
7B MLX teacher with best-of-K self-consistency (K=3) and a 6-gate… See the full description on the dataset page: https://huggingface.co/datasets/kakashi3lite/soulbox-cbt-therapy-dataset.darija-therapy-qa
Moroccan Darija Therapy Conversations Dataset
This dataset is entirely synthetic and contains no real patient information.
It is provided strictly for research, educational, and experimental purposes and must not be used for clinical, medical, diagnostic, or psychological decision-making.
Citation
If you use this dataset in your research, please cite:
@dataset{moroccan_darija_therapy_conversations,
title={Moroccan Darija Therapy Conversations},
author={Jamal… See the full description on the dataset page: https://huggingface.co/datasets/yibba/darija-therapy-qa.therapydata
Dataset Card for "therapydata"
More Information needed
finetuning_therapyeurope-who-estimated-number-of-children-needing-antiretroviral-therapy
Estimated number of children needing antiretroviral therapy based on WHO methods | Europe (WHO GHO)
🇪🇺 174 observations · 5 Europe countries · 1990–2024 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 174 observations of Estimated number of children needing antiretroviral therapy based on WHO methods data across 5 Europe countries, spanning 1990–2024, covering 1 distinct indicators.
About the source
Source: WHO Global… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-who-estimated-number-of-children-needing-antiretroviral-therapy.Therapydataset_formatted_807K
Dataset Card for "Therapydataset_formatted_807K"
More Information needed
dia-therapy-dataset
Dia Psychology Dataset
📚 Dataset Overview
The Dia Psychology Dataset is designed to train and evaluate AI models specializing in mental health conversations. It contains 9,850 structured question-answer pairs covering a broad spectrum of psychological topics. This dataset is ideal for building chatbots, virtual assistants, and AI models that provide empathetic and GenZ-friendly mental health support.
🛠 Data Collection & Processing
The dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/anupamaditya/dia-therapy-dataset.
