datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multilingual-Therapy-Dialogues
Dataset Summary
Multilingual Therapy Dialogues is a diverse and bilingual dataset consisting of paired dialogues between patients and therapists in both Persian and English.
Dataset Statistics
Number of samples: 7,179
English:
Average tokens per sentence: 101.30
Maximum tokens in a sentence: 939
Average characters per sentence: 567.85
Number of unique tokens: 32,968
Persian:
Average tokens per sentence: 100.06
Maximum tokens in a sentence: 1,413
Average… See the full description on the dataset page: https://huggingface.co/datasets/Algorithmic-Human-Development-Group/Multilingual-Therapy-Dialogues.Synthetic_Therapy_ConversationsTherapyDataset
nlpresearch
Dataset 1: MentalHealthDataset
Dataset 2: mentalhealth
Dataset 3: NLP Mental Health Conversations
Dataset 4: MentalHealthConversations
Dataset 5: Synthetic Therapy Conversations
Dataset 6: therapy-bot-data-10k
Dataset 7: Therapy-Alpaca
Dataset 8: merged_mental_health_dataset
therapy-conversations-combinedTherapy_sessions_datasetdia-therapy-dataset
Dia Psychology Dataset
📚 Dataset Overview
The Dia Psychology Dataset is designed to train and evaluate AI models specializing in mental health conversations. It contains 9,850 structured question-answer pairs covering a broad spectrum of psychological topics. This dataset is ideal for building chatbots, virtual assistants, and AI models that provide empathetic and GenZ-friendly mental health support.
🛠 Data Collection & Processing
The dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/anupamaditya/dia-therapy-dataset.ABX-CT-004_sequential_therapy_optimization_loss-v0.1ABX-CT-004 Sequential Therapy Optimization
Purpose
Detect when a planned Drug A then Drug B sequence loses effectiveness.
Core pattern
stress_index high
seq_gap rises and stays high
mono MICs stay below cutoffs during the early window
later failure_flag becomes 1
Files
data/train.csv
data/test.csv
scorer.py
Schema
Each row is one timepoint in a within strain series.
Required columns
row_id
series_id
timepoint_h
organism
strain_id
seq_drug_a
seq_drug_b
stress_index
seq_effectiveness_score… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ABX-CT-004_sequential_therapy_optimization_loss-v0.1.ACT_therapy_scriptsExamples patient conversations using ACT (Acceptance & Commitment therapy). Scripts come from 17 ACT books.
llm_therapy_dataTherapybotdatasettherapyTherapy_Diagnosisalternate_therapytherapy_emotions_ruSynthetic dataset of therapeutic data on russian language. Created by Claude Sonnet 4.6. Suitable for fine tuning multilabel classification models.
Therapytherapy-chatbot-files
