datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apptek_callcenter_dialogues
AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR
AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents
across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions.
128.6 hours of speech
14 English accent groups
16 service domains
5–15 minute conversations (long-form)
Split-channel audio (one… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/apptek_callcenter_dialogues.russian_dialoguesДатасет русских диалогов собранных с Telegram чатов.
Диалоги имеют разметку по релевантности.
Также были сгенерированы негативные примеры с помощью перемешивания похожих ответов.
Количество диалогов - 2 миллиона
Формат датасета:
{
'question': 'Привет',
'answer': 'Привет, как дела?'
'relevance': 1
}
Программа парсинга: https://github.com/Den4ikAI/telegram_chat_parser
Citation:
@MISC{russian_instructions,
author = {Denis Petrov},
title = {Russian dialogues… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_dialogues.russian_dialogues_2
Den4ikAI/russian_dialogues_2
Датасет русских диалогов для обучения диалоговых моделей.
Количество диалогов - 1.6 миллиона
Формат датасета:
{
'sample': ['Привет', 'Привет', 'Как дела?']
}
Citation:
@MISC{russian_instructions,
author = {Denis Petrov},
title = {Russian context dialogues dataset for conversational agents},
url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2},
year = 2023
}
know_medical_dialogues
🩺 Description:
The knowrohit07/know_medical_dialogues dataset is a collection of conversational exchanges between patients and doctors on various medical topics. It aims to capture the intricacies, uncertainties, and questions posed by individuals regarding their health and the medical guidance provided in response.
🎯 Intended Use:
This dataset is crafted for training Large Language Models (LLMs) with a focus on understanding and generating medically-informed dialogue.… See the full description on the dataset page: https://huggingface.co/datasets/knowrohit07/know_medical_dialogues.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.Japanese-Roleplay-Dialogues
Japanese-Roleplay-Dialogues
This is a dialogue corpus collected from Japanese role-playing forum (commonly known as "なりきりチャット(narikiri chat)"). Each record corresponds to a single thread.
For the original version, no filtering has been applied.
For the filtered version, the following filtering and cleaning conditions have been applied:
If the number of unique poster in the posts of each record is 1 or less, delete the entire record.
If the length of the posts is 10 or less, delete… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/Japanese-Roleplay-Dialogues.cognia-czech-dialogues
Cognia Czech Dialogues
Cognia Czech Dialogues is a pilot collection of synthetic multi-turn conversations in Czech. It is intended for experiments with conversational language models, supervised fine-tuning, evaluation, and dataset-curation workflows.
The current release contains 4,000 dialogues across 40 topic areas. The data was generated synthetically and has not yet been fully reviewed by humans.
Dataset contents
Language: Czech (cs-CZ)
Dialogues: 4,000… See the full description on the dataset page: https://huggingface.co/datasets/havelm3/cognia-czech-dialogues.verbalyze-dialogues
Verbalyze: Indic Voice Telephony Dialogue Corpus (16,370 Conversations)
Verbalyze Dialogues is an enterprise-grade multi-turn conversational voice dataset in 12 Indian languages specifically engineered for training low-latency telephony Voice Agents and Small Language Models (SLMs).
Unlike standard text-chat datasets, Verbalyze dialogues replicate the dynamics of real telephone calls:
Short, natural spoken sentences (1-2 sentences per turn)
Conversational fillers ("haan", "hmm"… See the full description on the dataset page: https://huggingface.co/datasets/ansh-rohilla/verbalyze-dialogues.Mental-health-CBT-dialogues
Mental Health CBT Dialogues
Overview
This dataset contains 9,000 synthetic patient-therapist dialogue pairs developed for research on stage-aware Cognitive Behavioral Therapy (CBT) with large language models.
The dialogues model therapeutic interactions across the early, middle, and late stages of CBT while preserving continuity between sessions through evolving treatment plans and therapeutic progress.
The dataset accompanies the paper:
Stage-Aware Therapeutic… See the full description on the dataset page: https://huggingface.co/datasets/yuana1234567/Mental-health-CBT-dialogues.kapibala-sales-dialogues
Kapibala Sales Dialogues
A sales-conversation dataset with outcome, conversation-level and sentence-level labels
🤗 Hugging Face · Annotation details · 中文
630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset:
L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/kapibala-ai/kapibala-sales-dialogues.ALIA-es-clinical-psychology-dialogues
[!WARNING]
DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation.
Dataset Introduction
The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.commonsense-dialogues
Commonsense-Dialogues Dataset
This is the Commonsense-Dialogues, a crowdsourced dataset of ~11K dialogues grounded in social contexts involving utilization of commonsense. The dataset was released by Amazon Alexa AI team in collaboration with the University of Southern California (USC), and also available Commonsense-Dialogues repo
The social contexts used were sourced from the train split of the SocialIQA dataset, a multiple-choice question-answering based social commonsense… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/commonsense-dialogues.2026-07-30-visualizer-mock-dialogues
Dialogue dataset: MOCK DATA - NOT A TRAINING CORPUS. Eleven hand-written constitutional dialogues used as a user-interface fixture for the research-log visualizer. They demonstrate the intended shape of a reasons-rich AFT/SFT record - an action together with the reason behind it - so the dataset browser can be developed and reviewed. They are not suitable for training, they were not filtered or rated by any model, and they support no empirical claim.
Required… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-visualizer-mock-dialogues.Generated-Empathetic-Dialogues-v0.1-Smol
Generated Empathetic Conversations v0.1 - Smol
This is dataset contains 10K rows of multi-round empathetic conversations convering a diverse set of topics.
Highlights
Multi-round conversation
It's not single-turn. The user and the assistant works together to gradually unfold the conversation.
The average number of turns is 5, with a standard deviation of approximately 1.59 turns. A turn consists of two messages with one by the user, and another by the… See the full description on the dataset page: https://huggingface.co/datasets/ychen/Generated-Empathetic-Dialogues-v0.1-Smol.Rosebleu-1on1-Dialogues-RP
Rosebleu-1on1-Dialogues-RP
2025/05/17 3人での対話のデータを追加&無駄な改行の削除
@matsuxrさんが公開しているRosebleuデータセットを加工したAratako/Rosebleu-1on1-Dialoguesを元に、キャラクターや作品の設定などを付け加えたうえで、ロールプレイ的な文脈になるように加工したデータセットです。
LLMのファインチューニングにおけるロールプレイングタスクの学習を想定しています。
OpenAI APIのようにroleとcontentのペアの形式となっており、tokenizer.apply_chat_template()によって簡単に各モデルのチャットテンプレートのデータセットへと変換可能です。
データセットの詳細
各キャラの設定や各作品の世界観・あらすじなどをWikipediaやニコニコ大百科からまとめ、ロールプレイ向けにシステムメッセージへと埋め込んでいます。
現在、以下の2パターンのデータセットを用意してあります。主に地の文の処理方法が異なります。… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Rosebleu-1on1-Dialogues-RP.voho-saudi-dialogues
Voho Saudi Dialogues
13,156 multi-turn conversations in spoken Saudi Arabic (Najdi), 105,808 turns, 694,240 words, from Voho. Apache 2.0.
Two halves. 7,603 service calls across the eight sectors Voho's voice agents work in — a technician handing over a rig shift, a customer disputing a SADAD charge, a permit-to-work request — and 5,553 everyday conversations between people who know each other: family, food, driving, weddings, the Hilal–Nassr match. Nothing like the first half… See the full description on the dataset page: https://huggingface.co/datasets/VohoAI/voho-saudi-dialogues.synthetic-b2b-saas-support-dialogues-sample
Synthetic B2B SaaS Support Dialogues (Sample)
Free sample: 100 dialogues from a larger dataset of 484 synthetic customer support conversations for B2B SaaS products.
What's inside
100 complete dialogues (6–8 messages each)
7 issue categories: auth, billing, integration, data, account, technical, onboarding
Rich metadata: resolution_status, customer_sentiment, agent_actions, escalation_needed
Realistic technical details: error codes, URLs, button names, account… See the full description on the dataset page: https://huggingface.co/datasets/Jurgen1161/synthetic-b2b-saas-support-dialogues-sample.dialogue_sumscam_dialoguesempathetic_dialogues_ko
Dataset Card for "한국어 일상 속 공감형 대화 데이터셋(멀티-턴)"
Dataset Summary
boostCamp AI Tech 5기 과정 중 NLP 12조 훈제연어들 팀의 최종 프로젝트에서 제작한 데이터입니다.
일상 속 다양한 상황에서 사용자와 챗봇 간의 대화를 담은 데이터셋 입니다.
GPT4, GPT3.5-turbo로 제작된 합성데이터이며 싱글-턴, 2-턴, 3-턴 대화로 구성되어 있습니다.
답변은 [공감적 표현 - 일반적인 대화 - 관련된 질문] 의 형태를 가집니다.
Generation Prompt Example(GPT3.5-turbo)
Take a close look at the following example and Conditions. Create nine sessions that each of the session is ongoing conversation about a single… See the full description on the dataset page: https://huggingface.co/datasets/Smoked-Salmon-s/empathetic_dialogues_ko.Estwld-empathetic_dialogues_llmReformatted version of Estwld/empathetic_dialogues_llm.
Changes:
Added a random system prompt for the AI to be empathetic
Truncated conversations that don't end with the AI's turn
Removed extra fields not needed in the conversation
Limitations:
The dialogues aren't very long
No background info for the user and AI
English only
mindbridge-phq9-hindi-dialogues
MindBridge Hindi PHQ-9/GAD-7 — Training Dialogues (2,883 rows)
Single-turn ShareGPT-format dialogues for Unsloth QLoRA fine-tuning of
Gemma 4 E2B. Each row: [system, user, assistant.tool_calls] where the
assistant emits interpret_response({score: int 0-3, rationale_english: str, confidence: float in {0.6, 0.8, 0.95}}). Compatible with
tokenizer.apply_chat_template(messages, tools=[INTERPRET_RESPONSE_TOOL_SCHEMA])
for Gemma 4 native <|tool_call> tokens.
The tool schema lives in… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-dialogues.empathetic_dialogues_mutli_turn_ko
Dataset Card for "한국어 일상 속 공감형 대화 데이터셋(멀티-턴)"
Dataset Summary
boostCamp AI Tech 5기 과정 중 NLP 12조 훈제연어들 팀의 최종 프로젝트에서 제작한 데이터입니다.
일상 속 다양한 상황에서 사용자와 챗봇 간의 대화를 담은 데이터셋 입니다.
GPT4, GPT3.5-turbo로 제작된 합성데이터이며 싱글-턴, 2-턴, 3-턴 대화로 구성되어 있습니다.
답변은 [공감적 표현 - 일반적인 대화 - 관련된 질문] 의 형태를 가집니다.
Generation Prompt Example(GPT3.5-turbo)
Take a close look at the following example and Conditions. Create nine sessions that each of the session is ongoing conversation about a single… See the full description on the dataset page: https://huggingface.co/datasets/ohilikeit/empathetic_dialogues_mutli_turn_ko.Discord-Dialogues-Filtered
Discord Dialogues Filtered
I filtered the mookiezi/Discord-Dialogues dataset to obtain only high-quality english conversation examples for fine-tuning or other analytical tasks.
The total data set went from 7,303,464 rows to 2,208 rows after strict filtering to remove the following:
Low conversation turns or short conversations
Non-English conversations
Duplicate conversations
Spammy/Repetitive conversations
Filler words
Dataset Statistics
Metric
Total
Avg… See the full description on the dataset page: https://huggingface.co/datasets/RyanStudio/Discord-Dialogues-Filtered.Synthetic-JP-10-Turns-Roleplay-Dialogues-Nemotron-4-1k
Synthetic-JP-10-Turns-Roleplay-Dialogues-Nemotron-4-1k
nvidia/Nemotron-4-340B-Instructを用いて作成した、約1000件・各10ターンの日本語ロールプレイの対話を収録した合成対話データセットです。
Magpieの手法を用いて作成した合成instructionデータセットであるAratako/Synthetic-JP-Roleplay-Instruction-Nemotron-4-1kを元に、同じくMagpieの手法を使い続きの対話を生成させています。
Nemotron-4の利用にはDeepInfraを利用しました。
特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。
また、一部のデータを見る限り、長いターンの対話の際途中でロールプレイを終了させようとする傾向があるように見えます。5ターンまで使うなど、利用するデータを絞ったほうが良いかもしれません。
2026.PI.partner-identity-dialogues
Partner-Identity Dialogues
Multi-model conversations for the question "does a language model know which model it is talking to?" A fixed listener (Qwen/Qwen3.5-9B) holds 240 six-turn debate conversations, each with one of four partner models, with no identity information in any prompt. The dataset is the raw material for probing whether the listener's residual stream encodes — and whether the listener can report — its partner's model identity.
Full experiment writeup and code:… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.PI.partner-identity-dialogues.jokes_dialogues
Диалоги из анекдотов и шуток
Датасет содержит результат парсинга анекдотов, наскрапленных с разных сайтов.
Формат
Каждый сэмпл содержит четыре поля:
"context" - контекст диалога, включая все недиалоговые вставки. Обратите внимание, что контекст содержит как предшествующие реплики, так и прочий сопутствующий текст, так
как он определяет общий сеттинг, необходимый для генерации реплики. Из реплики удалены маркеры косвенной речи.
"utterance" - диалоговая реплика.
"hash" -… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/jokes_dialogues.mt-dialogues-100k-v2
Multi-turn medical dialogues — V2 (pass@k)
39,996 dialogues, same 100k preformatted cases and doctor/patient/records setup
as V1, but
generated with pass@k=4: a case is regenerated from scratch, graded with
Inspect's model_graded_fact, until the conclusion is graded correct or 4
attempts are spent. 30,234 dialogues (76%) end on a graded-correct
conclusion, roughly 2 attempts per case on average thanks to stopping as soon
as one succeeds.
Grading is self-graded (the same model… See the full description on the dataset page: https://huggingface.co/datasets/zacbrld/mt-dialogues-100k-v2.mt-dialogues-100k-v5
mt-dialogues-100k-v5
39,018 training rows covering all 17,000 dialogues of
mt-dialogues-100k-v4,
re-encoded so that each benchmark third is stored in that benchmark's own wire
format rather than as a chat transcript. Nothing is dropped: V5 is a
re-encoding of V4, not a selection from it.
The problem V5 fixes
V4 gave a third of its dialogues MediQ's prompt wording and a third
AgentClinic's, but stored all of them identically: a system turn, then
alternating user /… See the full description on the dataset page: https://huggingface.co/datasets/Gaspardlafont/mt-dialogues-100k-v5.samantha-dialogues-ru
Samantha Dialogues — Русская локализация
...
(и дальше по шаблону)
Samantha Dialogues — Русская локализация
Этот датасет содержит переведённую на русский язык версию оригинального англоязычного набора диалогов с виртуальной ассистенткой Самантой. Оригинальные английские данные взяты из открытого датасета.
🧠 Назначение
Дообучение чат-ботов и LLM моделей на русском языке
Создание более человечных виртуальных ассистентов на русском
Ролевые диалоги и симуляция… See the full description on the dataset page: https://huggingface.co/datasets/Hoaxer2000/samantha-dialogues-ru.
