CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01apptek-com /apptek_callcenter_dialogues AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR AppTek Call-Center Dialogues is a long-form conversational speech dataset for automatic speech recognition (ASR), featuring diverse English accents across multiple service-oriented domains and designed to evaluate models on realistic call-center interactions. 128.6 hours of speech 14 English accent groups 16 service domains 5–15 minute conversations (long-form) Split-channel audio (one… See the full description on the dataset page: https://huggingface.co/datasets/apptek-com/apptek_callcenter_dialogues.audioautomatic-speech-recognition1K<n<10K37 likes2.5k downloads1mo agoHugging Face02Den4ikAI /russian_dialoguesДатасет русских диалогов собранных с Telegram чатов. Диалоги имеют разметку по релевантности. Также были сгенерированы негативные примеры с помощью перемешивания похожих ответов. Количество диалогов - 2 миллиона Формат датасета: { 'question': 'Привет', 'answer': 'Привет, как дела?' 'relevance': 1 } Программа парсинга: https://github.com/Den4ikAI/telegram_chat_parser Citation: @MISC{russian_instructions, author = {Denis Petrov}, title = {Russian dialogues… See the full description on the dataset page: https://huggingface.co/datasets/Den4ikAI/russian_dialogues.text1M<n<10M50 likes981 downloads4y agoHugging Face03Den4ikAI /russian_dialogues_2 Den4ikAI/russian_dialogues_2 Датасет русских диалогов для обучения диалоговых моделей. Количество диалогов - 1.6 миллиона Формат датасета: { 'sample': ['Привет', 'Привет', 'Как дела?'] } Citation: @MISC{russian_instructions, author = {Denis Petrov}, title = {Russian context dialogues dataset for conversational agents}, url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2}, year = 2023 } texttext-generation1M<n<10M18 likes597 downloads2y agoHugging Face04knowrohit07 /know_medical_dialogues 🩺 Description: The knowrohit07/know_medical_dialogues dataset is a collection of conversational exchanges between patients and doctors on various medical topics. It aims to capture the intricacies, uncertainties, and questions posed by individuals regarding their health and the medical guidance provided in response. 🎯 Intended Use: This dataset is crafted for training Large Language Models (LLMs) with a focus on understanding and generating medically-informed dialogue.… See the full description on the dataset page: https://huggingface.co/datasets/knowrohit07/know_medical_dialogues.textn<1K2 likes242 downloads3y agoHugging Face05Thomasgudan /kapibala-sales-dialogues Kapibala Sales Dialogues A sales-conversation dataset with outcome, conversation-level and sentence-level labels 🤗 Hugging Face · Annotation details · 中文 630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset: L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.tabulartext-classification10K<n<100K2 likes201 downloads8d agoHugging Face06OmniAICreator /Japanese-Roleplay-Dialogues Japanese-Roleplay-Dialogues This is a dialogue corpus collected from Japanese role-playing forum (commonly known as "なりきりチャット(narikiri chat)"). Each record corresponds to a single thread. For the original version, no filtering has been applied. For the filtered version, the following filtering and cleaning conditions have been applied: If the number of unique poster in the posts of each record is 1 or less, delete the entire record. If the length of the posts is 10 or less, delete… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/Japanese-Roleplay-Dialogues.texttext-generation10K<n<100K17 likes128 downloads2y agoHugging Face07havelm3 /cognia-czech-dialogues Cognia Czech Dialogues Cognia Czech Dialogues is a pilot collection of synthetic multi-turn conversations in Czech. It is intended for experiments with conversational language models, supervised fine-tuning, evaluation, and dataset-curation workflows. The current release contains 4,000 dialogues across 40 topic areas. The data was generated synthetically and has not yet been fully reviewed by humans. Dataset contents Language: Czech (cs-CZ) Dialogues: 4,000… See the full description on the dataset page: https://huggingface.co/datasets/havelm3/cognia-czech-dialogues.texttext-generation1K<n<10K0 likes109 downloads9d agoHugging Face08ansh-rohilla /verbalyze-dialogues Verbalyze: Indic Voice Telephony Dialogue Corpus (16,370 Conversations) Verbalyze Dialogues is an enterprise-grade multi-turn conversational voice dataset in 12 Indian languages specifically engineered for training low-latency telephony Voice Agents and Small Language Models (SLMs). Unlike standard text-chat datasets, Verbalyze dialogues replicate the dynamics of real telephone calls: Short, natural spoken sentences (1-2 sentences per turn) Conversational fillers ("haan", "hmm"… See the full description on the dataset page: https://huggingface.co/datasets/ansh-rohilla/verbalyze-dialogues.texttext-generation10K<n<100K1 likes108 downloads12d agoHugging Face09yuana1234567 /Mental-health-CBT-dialogues Mental Health CBT Dialogues Overview This dataset contains 9,000 synthetic patient-therapist dialogue pairs developed for research on stage-aware Cognitive Behavioral Therapy (CBT) with large language models. The dialogues model therapeutic interactions across the early, middle, and late stages of CBT while preserving continuity between sessions through evolving treatment plans and therapeutic progress. The dataset accompanies the paper: Stage-Aware Therapeutic… See the full description on the dataset page: https://huggingface.co/datasets/yuana1234567/Mental-health-CBT-dialogues.texttext-generation1K<n<10K4 likes95 downloads3mo agoHugging Face10kapibala-ai /kapibala-sales-dialogues Kapibala Sales Dialogues A sales-conversation dataset with outcome, conversation-level and sentence-level labels 🤗 Hugging Face · Annotation details · 中文 630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset: L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/kapibala-ai/kapibala-sales-dialogues.tabulartext-classification10K<n<100K0 likes74 downloads5d agoHugging Face11SINAI /ALIA-es-clinical-psychology-dialogues [!WARNING] DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation. Dataset Introduction The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.texttext-generationn<1K0 likes71 downloads3mo agoHugging Face12gretelai /commonsense-dialogues Commonsense-Dialogues Dataset This is the Commonsense-Dialogues, a crowdsourced dataset of ~11K dialogues grounded in social contexts involving utilization of commonsense. The dataset was released by Amazon Alexa AI team in collaboration with the University of Southern California (USC), and also available Commonsense-Dialogues repo The social contexts used were sourced from the train split of the SocialIQA dataset, a multiple-choice question-answering based social commonsense… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/commonsense-dialogues.texttext-classification10K<n<100K6 likes68 downloads2y agoHugging Face13dougalldeepmind /2026-07-30-visualizer-mock-dialogues Dialogue dataset: MOCK DATA - NOT A TRAINING CORPUS. Eleven hand-written constitutional dialogues used as a user-interface fixture for the research-log visualizer. They demonstrate the intended shape of a reasons-rich AFT/SFT record - an action together with the reason behind it - so the dataset browser can be developed and reviewed. They are not suitable for training, they were not filtered or rated by any model, and they support no empirical claim. Required… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-30-visualizer-mock-dialogues.textn<1K0 likes68 downloads1mo agoHugging Face14ychen /Generated-Empathetic-Dialogues-v0.1-Smol Generated Empathetic Conversations v0.1 - Smol This is dataset contains 10K rows of multi-round empathetic conversations convering a diverse set of topics. Highlights Multi-round conversation It's not single-turn. The user and the assistant works together to gradually unfold the conversation. The average number of turns is 5, with a standard deviation of approximately 1.59 turns. A turn consists of two messages with one by the user, and another by the… See the full description on the dataset page: https://huggingface.co/datasets/ychen/Generated-Empathetic-Dialogues-v0.1-Smol.texttext-generation10K<n<100K4 likes59 downloads2y agoHugging Face15Aratako /Rosebleu-1on1-Dialogues-RP Rosebleu-1on1-Dialogues-RP 2025/05/17 3人での対話のデータを追加&無駄な改行の削除 @matsuxrさんが公開しているRosebleuデータセットを加工したAratako/Rosebleu-1on1-Dialoguesを元に、キャラクターや作品の設定などを付け加えたうえで、ロールプレイ的な文脈になるように加工したデータセットです。 LLMのファインチューニングにおけるロールプレイングタスクの学習を想定しています。 OpenAI APIのようにroleとcontentのペアの形式となっており、tokenizer.apply_chat_template()によって簡単に各モデルのチャットテンプレートのデータセットへと変換可能です。 データセットの詳細 各キャラの設定や各作品の世界観・あらすじなどをWikipediaやニコニコ大百科からまとめ、ロールプレイ向けにシステムメッセージへと埋め込んでいます。 現在、以下の2パターンのデータセットを用意してあります。主に地の文の処理方法が異なります。… See the full description on the dataset page: https://huggingface.co/datasets/Aratako/Rosebleu-1on1-Dialogues-RP.texttext-generation1K<n<10K18 likes55 downloads2y agoHugging Face16VohoAI /voho-saudi-dialogues Voho Saudi Dialogues 13,156 multi-turn conversations in spoken Saudi Arabic (Najdi), 105,808 turns, 694,240 words, from Voho. Apache 2.0. Two halves. 7,603 service calls across the eight sectors Voho's voice agents work in — a technician handing over a rig shift, a customer disputing a SADAD charge, a permit-to-work request — and 5,553 everyday conversations between people who know each other: family, food, driving, weddings, the Hilal–Nassr match. Nothing like the first half… See the full description on the dataset page: https://huggingface.co/datasets/VohoAI/voho-saudi-dialogues.texttext-generation10K<n<100K0 likes53 downloads6d agoHugging Face17Jurgen1161 /synthetic-b2b-saas-support-dialogues-sample Synthetic B2B SaaS Support Dialogues (Sample) Free sample: 100 dialogues from a larger dataset of 484 synthetic customer support conversations for B2B SaaS products. What's inside 100 complete dialogues (6–8 messages each) 7 issue categories: auth, billing, integration, data, account, technical, onboarding Rich metadata: resolution_status, customer_sentiment, agent_actions, escalation_needed Realistic technical details: error codes, URLs, button names, account… See the full description on the dataset page: https://huggingface.co/datasets/Jurgen1161/synthetic-b2b-saas-support-dialogues-sample.texttext-generationn<1K0 likes50 downloads8d agoHugging Face18alierenak /dialogue_sumtext10K<n<100K0 likes41 downloads3y agoHugging Face19adamtc /scam_dialoguestexttext-classification1K<n<10K0 likes37 downloads2y agoHugging Face20Smoked-Salmon-s /empathetic_dialogues_ko Dataset Card for "한국어 일상 속 공감형 대화 데이터셋(멀티-턴)" Dataset Summary boostCamp AI Tech 5기 과정 중 NLP 12조 훈제연어들 팀의 최종 프로젝트에서 제작한 데이터입니다. 일상 속 다양한 상황에서 사용자와 챗봇 간의 대화를 담은 데이터셋 입니다. GPT4, GPT3.5-turbo로 제작된 합성데이터이며 싱글-턴, 2-턴, 3-턴 대화로 구성되어 있습니다. 답변은 [공감적 표현 - 일반적인 대화 - 관련된 질문] 의 형태를 가집니다. Generation Prompt Example(GPT3.5-turbo) Take a close look at the following example and Conditions. Create nine sessions that each of the session is ongoing conversation about a single… See the full description on the dataset page: https://huggingface.co/datasets/Smoked-Salmon-s/empathetic_dialogues_ko.texttext-generation10K<n<100K8 likes35 downloads3y agoHugging Face21agentlans /Estwld-empathetic_dialogues_llmReformatted version of Estwld/empathetic_dialogues_llm. Changes: Added a random system prompt for the AI to be empathetic Truncated conversations that don't end with the AI's turn Removed extra fields not needed in the conversation Limitations: The dialogues aren't very long No background info for the user and AI English only texttext-generation10K<n<100K0 likes34 downloads2y agoHugging Face22Huzayfah-Patel /mindbridge-phq9-hindi-dialogues MindBridge Hindi PHQ-9/GAD-7 — Training Dialogues (2,883 rows) Single-turn ShareGPT-format dialogues for Unsloth QLoRA fine-tuning of Gemma 4 E2B. Each row: [system, user, assistant.tool_calls] where the assistant emits interpret_response({score: int 0-3, rationale_english: str, confidence: float in {0.6, 0.8, 0.95}}). Compatible with tokenizer.apply_chat_template(messages, tools=[INTERPRET_RESPONSE_TOOL_SCHEMA]) for Gemma 4 native <|tool_call> tokens. The tool schema lives in… See the full description on the dataset page: https://huggingface.co/datasets/Huzayfah-Patel/mindbridge-phq9-hindi-dialogues.texttext-classification1K<n<10K0 likes33 downloads1mo agoHugging Face23ohilikeit /empathetic_dialogues_mutli_turn_ko Dataset Card for "한국어 일상 속 공감형 대화 데이터셋(멀티-턴)" Dataset Summary boostCamp AI Tech 5기 과정 중 NLP 12조 훈제연어들 팀의 최종 프로젝트에서 제작한 데이터입니다. 일상 속 다양한 상황에서 사용자와 챗봇 간의 대화를 담은 데이터셋 입니다. GPT4, GPT3.5-turbo로 제작된 합성데이터이며 싱글-턴, 2-턴, 3-턴 대화로 구성되어 있습니다. 답변은 [공감적 표현 - 일반적인 대화 - 관련된 질문] 의 형태를 가집니다. Generation Prompt Example(GPT3.5-turbo) Take a close look at the following example and Conditions. Create nine sessions that each of the session is ongoing conversation about a single… See the full description on the dataset page: https://huggingface.co/datasets/ohilikeit/empathetic_dialogues_mutli_turn_ko.texttext-generation10K<n<100K8 likes31 downloads3y agoHugging Face24RyanStudio /Discord-Dialogues-Filtered Discord Dialogues Filtered I filtered the mookiezi/Discord-Dialogues dataset to obtain only high-quality english conversation examples for fine-tuning or other analytical tasks. The total data set went from 7,303,464 rows to 2,208 rows after strict filtering to remove the following: Low conversation turns or short conversations Non-English conversations Duplicate conversations Spammy/Repetitive conversations Filler words Dataset Statistics Metric Total Avg… See the full description on the dataset page: https://huggingface.co/datasets/RyanStudio/Discord-Dialogues-Filtered.tabular1K<n<10K0 likes31 downloads1mo agoHugging Face25Aratako /Synthetic-JP-10-Turns-Roleplay-Dialogues-Nemotron-4-1k Synthetic-JP-10-Turns-Roleplay-Dialogues-Nemotron-4-1k nvidia/Nemotron-4-340B-Instructを用いて作成した、約1000件・各10ターンの日本語ロールプレイの対話を収録した合成対話データセットです。 Magpieの手法を用いて作成した合成instructionデータセットであるAratako/Synthetic-JP-Roleplay-Instruction-Nemotron-4-1kを元に、同じくMagpieの手法を使い続きの対話を生成させています。 Nemotron-4の利用にはDeepInfraを利用しました。 特に事後的なフィルタ処理は加えていないため、クオリティの低いレコードが含まれている可能性があります。ご注意ください。 また、一部のデータを見る限り、長いターンの対話の際途中でロールプレイを終了させようとする傾向があるように見えます。5ターンまで使うなど、利用するデータを絞ったほうが良いかもしれません。 texttext-generation1K<n<10K3 likes30 downloads2y agoHugging Face26siddharthmb /2026.PI.partner-identity-dialogues Partner-Identity Dialogues Multi-model conversations for the question "does a language model know which model it is talking to?" A fixed listener (Qwen/Qwen3.5-9B) holds 240 six-turn debate conversations, each with one of four partner models, with no identity information in any prompt. The dataset is the raw material for probing whether the listener's residual stream encodes — and whether the listener can report — its partner's model identity. Full experiment writeup and code:… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.PI.partner-identity-dialogues.tabulartext-classificationn<1K0 likes28 downloads3mo agoHugging Face27inkoziev /jokes_dialogues Диалоги из анекдотов и шуток Датасет содержит результат парсинга анекдотов, наскрапленных с разных сайтов. Формат Каждый сэмпл содержит четыре поля: "context" - контекст диалога, включая все недиалоговые вставки. Обратите внимание, что контекст содержит как предшествующие реплики, так и прочий сопутствующий текст, так как он определяет общий сеттинг, необходимый для генерации реплики. Из реплики удалены маркеры косвенной речи. "utterance" - диалоговая реплика. "hash" -… See the full description on the dataset page: https://huggingface.co/datasets/inkoziev/jokes_dialogues.text100K<n<1M5 likes24 downloads4y agoHugging Face28zacbrld /mt-dialogues-100k-v2 Multi-turn medical dialogues — V2 (pass@k) 39,996 dialogues, same 100k preformatted cases and doctor/patient/records setup as V1, but generated with pass@k=4: a case is regenerated from scratch, graded with Inspect's model_graded_fact, until the conclusion is graded correct or 4 attempts are spent. 30,234 dialogues (76%) end on a graded-correct conclusion, roughly 2 attempts per case on average thanks to stopping as soon as one succeeds. Grading is self-graded (the same model… See the full description on the dataset page: https://huggingface.co/datasets/zacbrld/mt-dialogues-100k-v2.tabulartext-generation10K<n<100K0 likes24 downloads1mo agoHugging Face29Gaspardlafont /mt-dialogues-100k-v5 mt-dialogues-100k-v5 39,018 training rows covering all 17,000 dialogues of mt-dialogues-100k-v4, re-encoded so that each benchmark third is stored in that benchmark's own wire format rather than as a chat transcript. Nothing is dropped: V5 is a re-encoding of V4, not a selection from it. The problem V5 fixes V4 gave a third of its dialogues MediQ's prompt wording and a third AgentClinic's, but stored all of them identically: a system turn, then alternating user /… See the full description on the dataset page: https://huggingface.co/datasets/Gaspardlafont/mt-dialogues-100k-v5.texttext-generation10K<n<100K0 likes22 downloads3d agoHugging Face30Hoaxer2000 /samantha-dialogues-ru Samantha Dialogues — Русская локализация ... (и дальше по шаблону) Samantha Dialogues — Русская локализация Этот датасет содержит переведённую на русский язык версию оригинального англоязычного набора диалогов с виртуальной ассистенткой Самантой. Оригинальные английские данные взяты из открытого датасета. 🧠 Назначение Дообучение чат-ботов и LLM моделей на русском языке Создание более человечных виртуальных ассистентов на русском Ролевые диалоги и симуляция… See the full description on the dataset page: https://huggingface.co/datasets/Hoaxer2000/samantha-dialogues-ru.text1K<n<10K3 likes19 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.