datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Discord-Dialogues
Discord-Dialogues is a large-scale dataset of anonymized Discord conversations from late spring to early fall 2025 for training and evaluating realistic conversational AI models in a ChatML-friendly format.
This dataset contains 7.3 million exchanges spread out over 16 million turns, with more than 139 million words.
Nomic Atlas Map
Features
Mixed single and multi-turn exchanges
Human-only dialogues (no bots)
Filtered for ToS and harmful contentLinks… See the full description on the dataset page: https://huggingface.co/datasets/mookiezi/Discord-Dialogues.Question-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.knowchat-multi-turn-dialogues
KnowChat: Multi-Turn Human-LLM Dialogues on Knowledge Tasks
KnowChat is a dataset of 705 multi-turn human-LLM conversations collected to validate the KnowSim user simulation framework. It pairs each conversation with pre/post knowledge assessments, self-reported survey ratings, and participant background information, enabling research on information calibration -- how well LLM assistants tailor responses to users with different knowledge levels.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/yjlee36/knowchat-multi-turn-dialogues.danish-tool-dialogues-v9
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
34,168
eval_seen_tools
698
eval_unseen_tools
768
eval_seen_sym
752… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v9.danish-tool-dialogues-v6
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,667
eval_seen_tools
722
eval_unseen_tools
779
933 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v6.danish-tool-dialogues-v7
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,160
eval_seen_tools
701
eval_unseen_tools
768
925 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v7.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.danish-tool-dialogues-v4
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,598
eval_seen_tools
762
eval_unseen_tools
772
932 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v4.danish-tool-dialogues-v5
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,596
eval_seen_tools
762
eval_unseen_tools
772
932 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v5.english-daily-dialogues-10k
English Daily Dialogues 10K
A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark.
Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.Discord-DialoguesThis is a clone of mookiezi/Discord-Dialogues.
Discord-Dialogues is a large-scale dataset of anonymized Discord conversations from late spring to early fall 2025 for training and evaluating realistic conversational AI models in a ChatML-friendly format.
This dataset contains 7.3 million exchanges spread out over 16 million turns, with more than 139 million words.
Features
Mixed single and multi-turn exchanges
Human-only dialogues (no bots)
Filtered for ToS and… See the full description on the dataset page: https://huggingface.co/datasets/aaronmoo12/Discord-Dialogues.danish-tool-dialogues-v8
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
34,168
eval_seen_tools
698
eval_unseen_tools
768
eval_seen_sym
752… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v8.danish-tool-dialogues-v3
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,354
eval_seen_tools
756
eval_unseen_tools
768
908 distinct tools; 58 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v3.danish-tool-dialogues-v1
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,138
eval_seen_tools
740
eval_unseen_tools
769
903 distinct tools; 57 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v1.danish-tool-dialogues-v2
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
17,354
eval_seen_tools
756
eval_unseen_tools
768
908 distinct tools; 58 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v2.empathetic_dialogues
Dataset Card for "empathetic_dialogues"
Dataset Summary
Dataset from Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset, but with parquet format.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 28.02 MB
Size of the generated dataset: 25.13 MB
Total amount of disk used:… See the full description on the dataset page: https://huggingface.co/datasets/brianist/empathetic_dialogues.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/kapibala-ai/kapibala-sales-dialogues.Discord-DialoguesThis is a clone of mookiezi/Discord-Dialogues.
Discord-Dialogues is a large-scale dataset of anonymized Discord conversations from late spring to early fall 2025 for training and evaluating realistic conversational AI models in a ChatML-friendly format.
This dataset contains 7.3 million exchanges spread out over 16 million turns, with more than 139 million words.
Features
Mixed single and multi-turn exchanges
Human-only dialogues (no bots)
Filtered for ToS and… See the full description on the dataset page: https://huggingface.co/datasets/mooaoeu/Discord-Dialogues.Hungarian-Dialogues-text
LLM-Generated Hungarian Conversations
This dataset contains structured Hungarian conversations generated with multiple large language model families for the study “Efficient ASR Training with Conversations that Never Happened.” Paper: arXiv link
Each model is provided as a separate Parquet file. The dataset contains the generated textual conversations and associated scenario and participant metadata; it does not contain synthesized audio.
Dataset structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/gedeonmate/Hungarian-Dialogues-text.empathetic_dialogues_with_special_tokensDiscord-Dialogues-Preprocessed-Luna-Protocol
Discord-Dialogues-Preprocessed-Luna-Protocol is a preprocessed fork of mookiezi/Discord-Dialogues, adapted for fine-tuning Qwen2.5-family models as part of the Luna Protocol project.
This dataset contains anonymized Discord conversations for training and evaluating realistic conversational AI models in a ChatML-friendly format. It is derived directly from mookiezi/Discord-Dialogues with two targeted preprocessing steps applied (see below) — the underlying conversations, filtering pipeline… See the full description on the dataset page: https://huggingface.co/datasets/fox3000foxy/Discord-Dialogues-Preprocessed-Luna-Protocol.empathetic_dialogues_frQuestion-Anchored-Tutoring-Dialogues-2k
Question-Anchored-Tutoring-Dialogues-2k
This dataset contains dialogues from math tutoring interventions recorded on Eedi.
Dataset Details
Dataset Description
Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data:
DQ-Question-Metadata: The question the student was working on that prompted the… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekh13/Question-Anchored-Tutoring-Dialogues-2k.Discord-Dialogues-Filtered
Discord Dialogues Filtered
I filtered the mookiezi/Discord-Dialogues dataset to obtain only high-quality english conversation examples for fine-tuning or other analytical tasks.
The total data set went from 7,303,464 rows to 2,208 rows after strict filtering to remove the following:
Low conversation turns or short conversations
Non-English conversations
Duplicate conversations
Spammy/Repetitive conversations
Filler words
Dataset Statistics
Metric
Total
Avg… See the full description on the dataset page: https://huggingface.co/datasets/RyanStudio/Discord-Dialogues-Filtered.2026.PI.partner-identity-dialogues
Partner-Identity Dialogues
Multi-model conversations for the question "does a language model know which model it is talking to?" A fixed listener (Qwen/Qwen3.5-9B) holds 240 six-turn debate conversations, each with one of four partner models, with no identity information in any prompt. The dataset is the raw material for probing whether the listener's residual stream encodes — and whether the listener can report — its partner's model identity.
Full experiment writeup and code:… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.PI.partner-identity-dialogues.roleplay_dialogues_extracted
Dialogues have not been extracted yet!
But characters do have been.
ru-stem-dialogues
Russian STEM Educational Dialogues
Описание
Синтетический датасет русскоязычных учебных диалогов по STEM-темам (математика, физика, химия,
биология, информатика, программирование, инженерия). Каждый диалог — реалистичное взаимодействие
между пользователем (школьник / студент / профессионал) и ассистентом.
Методология
Модель: Qwen/Qwen2.5-7B-Instruct (4-bit NF4 quantization, bitsandbytes)
Формат генерации: текстовый формат с разделителями… See the full description on the dataset page: https://huggingface.co/datasets/AtesiT/ru-stem-dialogues.mt-dialogues-100k-v2
Multi-turn medical dialogues — V2 (pass@k)
39,996 dialogues, same 100k preformatted cases and doctor/patient/records setup
as V1, but
generated with pass@k=4: a case is regenerated from scratch, graded with
Inspect's model_graded_fact, until the conclusion is graded correct or 4
attempts are spent. 30,234 dialogues (76%) end on a graded-correct
conclusion, roughly 2 attempts per case on average thanks to stopping as soon
as one succeeds.
Grading is self-graded (the same model… See the full description on the dataset page: https://huggingface.co/datasets/zacbrld/mt-dialogues-100k-v2.dstc2_dialogues_input_gpt2adaption-personal-finance-advice-dialogues
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-personal_finance_advice_dialogues
This dataset contains multi-turn conversational samples between users and an AI assistant focused on personal finance topics such as budgeting, investing, insurance, and taxes. Each entry follows a pattern where a user presents an initial scenario, provides an update with new constraints or events, and receives tailored financial advice that adapts… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-personal-finance-advice-dialogues.
