datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mental_health_counseling_conversations
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Amod/mental_health_counseling_conversations.opus-doctor-patient-conversations-all-human-diseases
Opus-4.8-High-Thinking generated Doctor-Patient Conversations for All Human Diseases
Covers every human disease listed on my previous work here: nisten/all-human-diseases
The dataset strictly used Opus 4.8 - High and was cleaned over 3 times via Opus 4.8, 4.7 and 4.6. Minor corrections were needed upon each pass mainly to bypass single word safety filters like i.e. monkeypox.
The main hallucination noticed during generation was that Opus would make up wrong PMID ( PubMed ID )… See the full description on the dataset page: https://huggingface.co/datasets/nisten/opus-doctor-patient-conversations-all-human-diseases.GLM-5.2-Conversation
GLM-5.2 · Conversation-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Token Count: 120M
Distribution:
Speaking domains:
•Greetings
•Customer Support
•Step by step explanations
•Motivational language
•Logical Questions
•Creative Writing
STEM:
•Algebra, calculus, quantum mechanics concepts
•Astromony and astrophysics
•Datascience and machine learning
•Biology
Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Conversation.Domofon-Cot-Conversations-700k
Domofon-Cot-Conversations-700k
Synthetic XML conversation data for training small language models on reasoning,
instruction following, XML formatting, and tool-use traces.
Repository: domofon/Domofon-Cot-Conversations-700k
What is inside
The dataset contains cleaned generated XML conversations from six families:
conv: multi-turn factual conversations with tool-use traces.
instruct: text-processing instructions, including deterministic count tool calls.
ds:… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Domofon-Cot-Conversations-700k.pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset
An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub.
The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language.
The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.scraped-chatgpt-conversations
Dataset Card for Dataset Name
Dataset Summary
scraped-chatgpt-conversations contains ~100k conversations between a user and chatgpt that were shared online through reddit, twitter, or sharegpt. For sharegpt, the conversations were directly scraped from the website. For reddit and twitter, images were downloaded from submissions, segmented, and run through an OCR pipeline to obtain a conversation list. For information on how the each json file is structured, please see… See the full description on the dataset page: https://huggingface.co/datasets/ar852/scraped-chatgpt-conversations.Mental-Health-Conversations
Dataset Card
This dataset consists of around 99k rows of mental health conversations. It is a cleaned version of "jerryjalapeno/nart-100k-synthetic".
Source
jerryjalapeno/nart-100k-synthetic
conversation-calendar
Overview
This dataset was generated for evaluation purposes, focusing on compiling an individual’s information with questions and answers.
Use of ChatGPT-4o: Used for generating human-like text and structured calendar data.
Format: For easy processing, data is stored in JSON (calendar data) and TXT (conversation data) formats.
Persona: The dataset is centered around an individual named ‘Alex’.
Event Generation: Calendar events are designed to be realistic rather than randomly… See the full description on the dataset page: https://huggingface.co/datasets/asu-kim/conversation-calendar.Bambara-dataset_conversation
license: apache-2.0
language:
- fr
- bm
tags:
- bambara
- bamanankan
- instruction-tuning
- llm-alignment
- african-languages
- low-resource-nlp
- conversational-ai
task_categories:
- text-generation
- conditional-text-generation
size_categories:
- 10K<n<50K
pretty_name: Bambara Instruction Tuning Corpus (FR-BM)
🌍 Bambara Instruction Tuning Corpus (FR-BM)
🚀 Overview & Vision
Welcome to the Bambara Instruction Tuning Corpus… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara-dataset_conversation.bengali-medical-triage-conversations
🩺 Bengali Medical Triage Conversations: Multilingual Clinical Dialogue & Screening Dataset
Bengali-Medical-Triage-Conversations is a medically audited, multi-turn clinical conversation dataset designed for fine-tuning and evaluating healthcare AI assistants and diagnostic triage agents in Bengali (বাংলা), Banglish (Phonetic Romanized Bengali), and English.
It is specifically tailored to address acute tropical, infectious, and primary care presentations prevalent in Bangladesh… See the full description on the dataset page: https://huggingface.co/datasets/Irtisum/bengali-medical-triage-conversations.socratic-method-conversations
Socratic Method Conversations Dataset
Overview
This dataset contains 5,000 question-answer pairs that demonstrate the Socratic method of teaching through guided questioning. The dataset has been carefully cleaned to remove all romantic and potentially inappropriate content, making it suitable for educational applications.
Dataset Description
The Socratic method is a form of inquiry and discussion between individuals, based on asking and answering questions to… See the full description on the dataset page: https://huggingface.co/datasets/sanjaypantdsd/socratic-method-conversations.GLM-5.2-Conversation
GLM-5.2 · Conversation-50000x
50,000x traces distilled from GLM-5.2 on High reasoning
Token Count: 120M
Distribution:
Speaking domains:
•Greetings
•Customer Support
•Step by step explanations
•Motivational language
•Logical Questions
•Creative Writing
STEM:
•Algebra, calculus, quantum mechanics concepts
•Astromony and astrophysics
•Datascience and machine learning
•Biology
Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/GLM-5.2-Conversation.FStarDataset-V2-Conversation
F* Proof Completion Dataset (Chat Format)
This dataset is a preprocessed version of microsoft/FStarDataSet-V2. It has been reformatted into a chat-style JSONL structure for supervised fine-tuning of language models on F* function synthesis and proof completion.
Dataset Structure
The dataset consists of three splits:
fstar_train.jsonl
fstar_validation.jsonl
fstar_test.jsonl
Each line in these files is a JSON object with the following schema (where the keys correspond to… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/FStarDataset-V2-Conversation.mental_health_counseling_conversations
Amod/mental_health_counseling_conversations
This data is cloned from https://huggingface.co/datasets/Amod/mental_health_counseling_conversations
Dataset Summary
This dataset is a collection of questions and answers sourced from two online counseling and therapy platforms. The questions cover a wide range of mental health topics, and the answers are provided by qualified psychologists. The dataset is intended to be used for fine-tuning language models to improve their… See the full description on the dataset page: https://huggingface.co/datasets/MaggiePai/mental_health_counseling_conversations.actuarial-gpt-conversations
👋 Connect with me on LinkedIn!
Manuel Caccone - Actuarial Data Scientist & Open Source Educator
Let's discuss actuarial science, AI, and open source projects!
📊 ActuarialGPT Conversations Dataset
Precision Mathematical Conversations for Insurance Intelligence
🎯 Quick Facts
Feature
Description
Domain
Actuarial Science, Insurance Analytics, Risk Management
Language
English (Technical/Expert Level)… See the full description on the dataset page: https://huggingface.co/datasets/manuelcaccone/actuarial-gpt-conversations.reasoning_conversations_advanced_1m
💻 Reasoning Conversations Advanced 1M
📖 Dataset Summary
Reasoning Conversations Advanced 1M is a massive-scale, synthetic dataset specifically engineered to improve the algorithmic reasoning and problem-solving capabilities of Large Language Models (LLMs). Featuring 1,000,000 unique coding samples, this dataset spans multiple programming languages (Python, JS, C++, etc.) and focuses on logic-heavy development tasks.
A key feature of this dataset is its Adaptive… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning_conversations_advanced_1m.afghanistan-post-2021-pashto-conversation-3x
🇦🇫 Afghanistan Post-2021 Pashto Conversation 3X
nassimjp/afghanistan-post-2021-pashto-conversation-3x
A Pashto conversational dataset focused on Afghanistan after 2021, designed for training and evaluating Pashto language models on multi-turn dialogue, answer diversity, contextual follow-up questions, and conversational continuity.
📌 Overview
This dataset is designed as a conversational extension of the Afghanistan Post-2021 Pashto Dataset.
Instead of providing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/afghanistan-post-2021-pashto-conversation-3x.comte-monte-cristo-conversations
Edmond Dantès Conversation Dataset
This dataset contains synthetic conversational data and source citations for fine-tuning language models to embody the character of Edmond Dantès from Alexandre Dumas' classic novel "Le Comte de Monte-Cristo" (The Count of Monte Cristo). The conversations are in formal 19th-century French, maintaining the literary style and personality of the protagonist.
The dataset includes two configurations:
conversations (default): 4,091 instruction-response… See the full description on the dataset page: https://huggingface.co/datasets/1ou2/comte-monte-cristo-conversations.llm-jp-chatbot-arena-conversations
LLM-jp Chatbot Arena Conversations Dataset
This dataset contains approximately 1,000 conversations with pairwise human preferences, most of which are in Japanese.
The data was collected during the trial phase of the LLM-jp Chatbot Arena (January–February 2025), where users compared responses from two different models in a head-to-head format.
Each sample includes a question ID, the names of the two models, their conversation transcripts, the user's vote, an anonymized user ID, a… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-chatbot-arena-conversations.mental_health_counseling_conversations_rated
Dataset Card for Mental Health Counseling Conversations Rated
This dataset extends the existing dataset Mental Health Counseling Conversations and adds ratings for the responses.
Dataset Details
This dataset is an extension for the dataset Mental Health Counseling Conversations.
It adds ratings for the responses generated by four different LLMs. The responses are rated across the following dimensions:
empathy
appropriateness
relevance
The following four LLMs are used… See the full description on the dataset page: https://huggingface.co/datasets/tcabanski/mental_health_counseling_conversations_rated.synthetic-neurology-conversations
Synthetic Neurology Conversations
Summary.This dataset augments questions from KryptoniteCrown/synthetic-neurology-QA-dataset with a compact two-step follow-up conversation generated by moonshotai/Kimi-K2-Instruct:
model answers the original question,
model asks a follow-up question (to deepen/clarify),
model answers its follow-up.
Shared by the OpenMed Community to help improve medical models globally.
Not medical advice. Research/education only; not for clinical… See the full description on the dataset page: https://huggingface.co/datasets/openmed-community/synthetic-neurology-conversations.JobCCC-Conversational-Job-Recommendation-Bangladesh
JobCCC: A Conversational Code-Mixed Corpus for Job Recommendation in Bangladesh
Dataset Creators
Authors: Md. Arman Hossain, Mubashir Jawad, Fariha Khandakar Moon, and Sonia Binte Siraj
Supervisor: Dr. Nafis Sadeq
Institution: Department of Computer Science & Engineering, East West University
Dataset Summary
JobCCC (Conversational Code-Mixed Corpus) is a multi-turn conversational benchmark and job recommendation dataset tailored for the… See the full description on the dataset page: https://huggingface.co/datasets/Armans33115/JobCCC-Conversational-Job-Recommendation-Bangladesh.mental_health_counseling_conversations-kk
🧠 Dataset Card — Kazakh Mental Health Counseling Conversations
Please ❤️Like❤️ this repo and if you like (and/or use) my work, thank you!
📚 Dataset Description
This dataset contains translated mental health counseling dialogues from English to Kazakh (қазақ тілі). The original source is the mental_health_counseling_conversations dataset by Amod, which has been cleaned and translated using Google Gemini API.
🗃 Dataset Structure
Format: CSV
Fields:… See the full description on the dataset page: https://huggingface.co/datasets/Eraly-ml/mental_health_counseling_conversations-kk.cat_conversations_jpSaudi-Arabic-Alzheimers-Conversational-Dataset-Parameterized
Saudi Arabic Alzheimer's Patient QA Dataset (Conversational)
Overview
This dataset contains parameterized question-answer pairs designed for conversational AI assistants supporting Alzheimer's patients. The questions are written in the Saudi Arabic dialect and cover common memory-related interactions.
Features
Saudi Arabic dialect
Parameterized answers
Alzheimer's memory support
Conversational QA
RAG-ready
Language
Arabic (Saudi… See the full description on the dataset page: https://huggingface.co/datasets/ShahadAljohani/Saudi-Arabic-Alzheimers-Conversational-Dataset-Parameterized.wiki-facts-conversations
Dataset Card for Ukrainian Wiki Facts Dialogs
Dataset Description
Dataset Summary
This dataset is a processed version of a cleaned Wikipedia text. Articles are summarized using Lapa LLM to provide key information about the topic asked. As an output, it contains summaries and dialogs, consisting of the following format:
>> Населення Американського Самоа
Чисельність населення країни становить 54,3 тисячі осіб. Природний приріст населення негативний, народжуваність становить… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/wiki-facts-conversations.neuroforge-conversation-dataset
NeuroForge Conversation Dataset
Overview
NeuroForge Conversation Dataset is an open-source dialogue dataset designed for developing real-time conversational AI systems.
The dataset contains structured human-assistant conversations that help language models learn natural communication, context understanding, and assistant behavior.
The goal is to support the development of lightweight and intelligent Small Language Models (SLMs).
Purpose
This… See the full description on the dataset page: https://huggingface.co/datasets/neuroforge-labs/neuroforge-conversation-dataset.harry_potter_conversational
Dataset: Harry potter conversational text corpus
Dataset Details
This corpus contains conversational data in text format
Usage
text classification
token classification
question answering
Language
en
License
apache 2.0
CONVERSATIONS_WITH_ANOTHER_LIFE_FORMThe research and factual part of this work would be incomplete without paying certain attention to the contacts of the Volga Group for the Study of UFOs with an unidentified source (or sources) of intelligent information. These contacts were carried out by us from the end of 1993 to 1997, i.e., over a period of five years. During this time, a rather extraordinary material of an intellectual nature has been accumulated, which needs to be deeply understood and, if possible, to draw certain… See the full description on the dataset page: https://huggingface.co/datasets/AndreySokolov01/CONVERSATIONS_WITH_ANOTHER_LIFE_FORM.
