CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Den4ikAI /russian_dialogues_2 Den4ikAI/russian_dialogues_2 Датасет русских диалогов для обучения диалоговых моделей. Количество диалогов - 1.6 миллиона Формат датасета: { 'sample': ['Привет', 'Привет', 'Как дела?'] } Citation: @MISC{russian_instructions, author = {Denis Petrov}, title = {Russian context dialogues dataset for conversational agents}, url = {https://huggingface.co/datasets/Den4ikAI/russian_dialogues_2}, year = 2023 } texttext-generation1M<n<10M18 likes597 downloads2y agoHugging Face023nesdeniz /turkish-daily-dialogues-5k Turkish Daily Dialogues 5K Exactly 5,000 synthetic, multi-turn Turkish conversations covering ordinary daily-life situations. The corpus is designed as a small, auditable baseline for dialogue modelling, instruction-format experiments, augmentation research, and Turkish-language evaluation—not as a substitute for conversations written by real people. Provenance in one sentence: the Turkish source scenario library was drafted with AI assistance specifically for this project… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/turkish-daily-dialogues-5k.texttext-generation1K<n<10K2 likes510 downloads2mo agoHugging Face03Eedi /Question-Anchored-Tutoring-Dialogues-2k Question-Anchored-Tutoring-Dialogues-2k This dataset contains dialogues from math tutoring interventions recorded on Eedi. Dataset Details Dataset Description Each dialogue represents a chat-based conversation between a tutor and a student prompted by the student requesting assistance while working on a lesson. Dialogues are accompanied with 2 sources of meta-data: DQ-Question-Metadata: The question the student was working on that prompted the tutoring… See the full description on the dataset page: https://huggingface.co/datasets/Eedi/Question-Anchored-Tutoring-Dialogues-2k.tabulartext-generation10K<n<100K10 likes488 downloads7mo agoHugging Face04yjlee36 /knowchat-multi-turn-dialogues KnowChat: Multi-Turn Human-LLM Dialogues on Knowledge Tasks KnowChat is a dataset of 705 multi-turn human-LLM conversations collected to validate the KnowSim user simulation framework. It pairs each conversation with pre/post knowledge assessments, self-reported survey ratings, and participant background information, enabling research on information calibration -- how well LLM assistants tailor responses to users with different knowledge levels. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/yjlee36/knowchat-multi-turn-dialogues.tabularquestion-answeringn<1K3 likes432 downloads1mo agoHugging Face05Estwld /empathetic_dialogues_llm Empathetic Dialogues for LLM  This repository contains a reformatted version of the Empathetic Dialogues dataset, tailored for seamless integration with Language Model (LLM) training and inference. The original dataset's format posed challenges for direct application in LLM tasks, prompting us to restructure and clean the data.  Data Restructuring  We have implemented the following changes to enhance the dataset's usability:  Merged dialogues with the same conv_id… See the full description on the dataset page: https://huggingface.co/datasets/Estwld/empathetic_dialogues_llm.texttext-generation10K<n<100K33 likes364 downloads2y agoHugging Face06jensjepsen /danish-tool-dialogues-v9 danish-tool-dialogues-v1 Danish multi-turn tool-use conversations with reasoning, translated from the Glaive subset of Nanbeige/ToolMind (Apache-2.0) by scripts/translate_toolmind_da.py. Complements danish-tool-calls-v1, which is single-turn and synthetic. Here the conversations run several turns, tool results are fed back, and the assistant reasons before calling. split rows train 34,168 eval_seen_tools 698 eval_unseen_tools 768 eval_seen_sym 752… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v9.tabulartext-generation100K<n<1M0 likes297 downloads17d agoHugging Face07jensjepsen /danish-tool-dialogues-v6 danish-tool-dialogues-v1 Danish multi-turn tool-use conversations with reasoning, translated from the Glaive subset of Nanbeige/ToolMind (Apache-2.0) by scripts/translate_toolmind_da.py. Complements danish-tool-calls-v1, which is single-turn and synthetic. Here the conversations run several turns, tool results are fed back, and the assistant reasons before calling. split rows train 17,667 eval_seen_tools 722 eval_unseen_tools 779 933 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v6.tabulartext-generation10K<n<100K0 likes219 downloads18d agoHugging Face08allenai /sdsd-dialogues Self Directed Synthetic Dialogues (SDSD) v0 This dataset is an experiment in procedurally generating synthetic dialogues between two language models. For each dialogue, one model, acting as a "user" generates a plan based on a topic, subtopic, and goal for a conversation. Next, this model attempts to act on this plan and generating synthetic data. Along with the plan is a principle which the model, in some successful cases, tries to cause the model to violate the principle resulting… See the full description on the dataset page: https://huggingface.co/datasets/allenai/sdsd-dialogues.texttext-generation100K<n<1M19 likes207 downloads2y agoHugging Face09jensjepsen /danish-tool-dialogues-v7 danish-tool-dialogues-v1 Danish multi-turn tool-use conversations with reasoning, translated from the Glaive subset of Nanbeige/ToolMind (Apache-2.0) by scripts/translate_toolmind_da.py. Complements danish-tool-calls-v1, which is single-turn and synthetic. Here the conversations run several turns, tool results are fed back, and the assistant reasons before calling. split rows train 17,160 eval_seen_tools 701 eval_unseen_tools 768 925 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v7.tabulartext-generation10K<n<100K0 likes202 downloads18d agoHugging Face10Thomasgudan /kapibala-sales-dialogues Kapibala Sales Dialogues A sales-conversation dataset with outcome, conversation-level and sentence-level labels 🤗 Hugging Face · Annotation details · 中文 630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset: L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/Thomasgudan/kapibala-sales-dialogues.tabulartext-classification10K<n<100K2 likes201 downloads8d agoHugging Face11jensjepsen /danish-tool-dialogues-v4 danish-tool-dialogues-v1 Danish multi-turn tool-use conversations with reasoning, translated from the Glaive subset of Nanbeige/ToolMind (Apache-2.0) by scripts/translate_toolmind_da.py. Complements danish-tool-calls-v1, which is single-turn and synthetic. Here the conversations run several turns, tool results are fed back, and the assistant reasons before calling. split rows train 17,598 eval_seen_tools 762 eval_unseen_tools 772 932 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v4.tabulartext-generation10K<n<100K0 likes200 downloads20d agoHugging Face12jensjepsen /danish-tool-dialogues-v5 danish-tool-dialogues-v1 Danish multi-turn tool-use conversations with reasoning, translated from the Glaive subset of Nanbeige/ToolMind (Apache-2.0) by scripts/translate_toolmind_da.py. Complements danish-tool-calls-v1, which is single-turn and synthetic. Here the conversations run several turns, tool results are fed back, and the assistant reasons before calling. split rows train 17,596 eval_seen_tools 762 eval_unseen_tools 772 932 distinct tools; 59 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v5.tabulartext-generation10K<n<100K0 likes152 downloads20d agoHugging Face13OmniAICreator /Japanese-Roleplay-Dialogues Japanese-Roleplay-Dialogues This is a dialogue corpus collected from Japanese role-playing forum (commonly known as "なりきりチャット(narikiri chat)"). Each record corresponds to a single thread. For the original version, no filtering has been applied. For the filtered version, the following filtering and cleaning conditions have been applied: If the number of unique poster in the posts of each record is 1 or less, delete the entire record. If the length of the posts is 10 or less, delete… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/Japanese-Roleplay-Dialogues.texttext-generation10K<n<100K17 likes128 downloads2y agoHugging Face143nesdeniz /english-daily-dialogues-10k English Daily Dialogues 10K A general-purpose, open dataset of 10,000 synthetic multi-turn English conversations spanning ten everyday-life domains. Built as a clean NLP resource for dialogue modeling, response generation, intent understanding, and conversational evaluation. This is a general language resource — not a safety or security benchmark. Curated by Enes Deniz (ORCID 0009-0006-9491-3565), Co-Founder at AltaySec. It is the English companion to the Turkish Daily Dialogues… See the full description on the dataset page: https://huggingface.co/datasets/3nesdeniz/english-daily-dialogues-10k.tabulartext-generation10K<n<100K2 likes127 downloads2mo agoHugging Face15Mykes /rus_med_dialogues Russian-language dataset of 2282 patient conversations in a medical bot. The training sample includes 2053 conversations; The test sample includes 229 conversations; Feature characteristics: topic - medical topic context - user-ai message history user_question - last user question assistant_answer - ai answer according the context and topic prompt - ready prompt for fincetuning instruct model (adapted for using with unsloth… See the full description on the dataset page: https://huggingface.co/datasets/Mykes/rus_med_dialogues.textquestion-answering1K<n<10K5 likes120 downloads2y agoHugging Face16jensjepsen /danish-tool-dialogues-v8 danish-tool-dialogues-v1 Danish multi-turn tool-use conversations with reasoning, translated from the Glaive subset of Nanbeige/ToolMind (Apache-2.0) by scripts/translate_toolmind_da.py. Complements danish-tool-calls-v1, which is single-turn and synthetic. Here the conversations run several turns, tool results are fed back, and the assistant reasons before calling. split rows train 34,168 eval_seen_tools 698 eval_unseen_tools 768 eval_seen_sym 752… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v8.tabulartext-generation100K<n<1M0 likes114 downloads17d agoHugging Face17jensjepsen /danish-tool-dialogues-v3 danish-tool-dialogues-v1 Danish multi-turn tool-use conversations with reasoning, translated from the Glaive subset of Nanbeige/ToolMind (Apache-2.0) by scripts/translate_toolmind_da.py. Complements danish-tool-calls-v1, which is single-turn and synthetic. Here the conversations run several turns, tool results are fed back, and the assistant reasons before calling. split rows train 17,354 eval_seen_tools 756 eval_unseen_tools 768 908 distinct tools; 58 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v3.tabulartext-generation10K<n<100K0 likes111 downloads21d agoHugging Face18havelm3 /cognia-czech-dialogues Cognia Czech Dialogues Cognia Czech Dialogues is a pilot collection of synthetic multi-turn conversations in Czech. It is intended for experiments with conversational language models, supervised fine-tuning, evaluation, and dataset-curation workflows. The current release contains 4,000 dialogues across 40 topic areas. The data was generated synthetically and has not yet been fully reviewed by humans. Dataset contents Language: Czech (cs-CZ) Dialogues: 4,000… See the full description on the dataset page: https://huggingface.co/datasets/havelm3/cognia-czech-dialogues.texttext-generation1K<n<10K0 likes109 downloads9d agoHugging Face19ansh-rohilla /verbalyze-dialogues Verbalyze: Indic Voice Telephony Dialogue Corpus (16,370 Conversations) Verbalyze Dialogues is an enterprise-grade multi-turn conversational voice dataset in 12 Indian languages specifically engineered for training low-latency telephony Voice Agents and Small Language Models (SLMs). Unlike standard text-chat datasets, Verbalyze dialogues replicate the dynamics of real telephone calls: Short, natural spoken sentences (1-2 sentences per turn) Conversational fillers ("haan", "hmm"… See the full description on the dataset page: https://huggingface.co/datasets/ansh-rohilla/verbalyze-dialogues.texttext-generation10K<n<100K1 likes108 downloads12d agoHugging Face20yuana1234567 /Mental-health-CBT-dialogues Mental Health CBT Dialogues Overview This dataset contains 9,000 synthetic patient-therapist dialogue pairs developed for research on stage-aware Cognitive Behavioral Therapy (CBT) with large language models. The dialogues model therapeutic interactions across the early, middle, and late stages of CBT while preserving continuity between sessions through evolving treatment plans and therapeutic progress. The dataset accompanies the paper: Stage-Aware Therapeutic… See the full description on the dataset page: https://huggingface.co/datasets/yuana1234567/Mental-health-CBT-dialogues.texttext-generation1K<n<10K4 likes95 downloads3mo agoHugging Face21Nerthus-Project /Generated_OE_Gregory_Dialogues_Text_and_Evaluation Generated Old English Gregory's Dialogues (variatio) A complete, machine-generated Old English variatio of the Old English Dialogues of Gregory the Great (Waerferth's translation), produced on 19 July 2026, together with the full generation and evaluation apparatus: prompt, constraint lexicon scripts, validator, dependency parses, word embeddings, and all quantitative evaluation results. The project is described in: Martin Arista, J., & Nunez, M. Evaluating Generated Old… See the full description on the dataset page: https://huggingface.co/datasets/Nerthus-Project/Generated_OE_Gregory_Dialogues_Text_and_Evaluation.text-generation1K<n<10K0 likes90 downloads17d agoHugging Face22jensjepsen /danish-tool-dialogues-v1 danish-tool-dialogues-v1 Danish multi-turn tool-use conversations with reasoning, translated from the Glaive subset of Nanbeige/ToolMind (Apache-2.0) by scripts/translate_toolmind_da.py. Complements danish-tool-calls-v1, which is single-turn and synthetic. Here the conversations run several turns, tool results are fed back, and the assistant reasons before calling. split rows train 17,138 eval_seen_tools 740 eval_unseen_tools 769 903 distinct tools; 57 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v1.tabulartext-generation10K<n<100K0 likes90 downloads23d agoHugging Face23jensjepsen /danish-tool-dialogues-v2 danish-tool-dialogues-v1 Danish multi-turn tool-use conversations with reasoning, translated from the Glaive subset of Nanbeige/ToolMind (Apache-2.0) by scripts/translate_toolmind_da.py. Complements danish-tool-calls-v1, which is single-turn and synthetic. Here the conversations run several turns, tool results are fed back, and the assistant reasons before calling. split rows train 17,354 eval_seen_tools 756 eval_unseen_tools 768 908 distinct tools; 58 names… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v2.tabulartext-generation10K<n<100K0 likes90 downloads22d agoHugging Face24brianist /empathetic_dialogues Dataset Card for "empathetic_dialogues" Dataset Summary Dataset from Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset, but with parquet format. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances default Size of downloaded dataset files: 28.02 MB Size of the generated dataset: 25.13 MB Total amount of disk used:… See the full description on the dataset page: https://huggingface.co/datasets/brianist/empathetic_dialogues.tabulartext-generation10K<n<100K1 likes78 downloads7mo agoHugging Face25kukunechka /russian-everyday-dialogues Russian Everyday Dialogues / Русские повседневные диалоги A dataset of 20 natural conversational Russian dialogues covering everyday situations. Создан носителем языка с 18-летним опытом работы в международной среде. Dataset Description This dataset contains short, natural Russian dialogues typical of everyday urban life in Russia. Each example reflects authentic spoken language patterns, not formal or literary Russian. Situations covered 🛒 Магазин / Shopping… See the full description on the dataset page: https://huggingface.co/datasets/kukunechka/russian-everyday-dialogues.text-generationn<1K0 likes75 downloads5mo agoHugging Face26kapibala-ai /kapibala-sales-dialogues Kapibala Sales Dialogues A sales-conversation dataset with outcome, conversation-level and sentence-level labels 🤗 Hugging Face · Annotation details · 中文 630 synthetic sales conversations (11,688 messages, Chinese and English, five domains) between an LLM-simulated customer and an AI salesperson. Every conversation carries three layers of labels, each produced by a single method across the whole dataset: L1 — outcome. Did the customer buy, agree to a next step, stay undecided… See the full description on the dataset page: https://huggingface.co/datasets/kapibala-ai/kapibala-sales-dialogues.tabulartext-classification10K<n<100K0 likes74 downloads5d agoHugging Face27SINAI /ALIA-es-clinical-psychology-dialogues [!WARNING] DISCLAIMER: This dataset is not clinically validated. It is a research proof-of-concept. It should not be used as clinical truth or as a replacement for qualified human professional consultation. Dataset Introduction The ALIA Spanish Clinical Psychology Dialogues Corpus is a curated conversational instruction-tuning resource in Spanish created under the ALIA project. It was designed to train and evaluate language models in empathetic therapeutic dialogue and… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-clinical-psychology-dialogues.texttext-generationn<1K0 likes71 downloads3mo agoHugging Face28Mykes /rus_med_dialogues_qa Russian-language dataset of 3941 patient conversations with a medical bot in QA manner. The training sample includes 3546 conversations; The test sample includes 395 conversations; Feature characteristics: topic - medical topic user_question - last user question assistant_answer - ai answer according the context and topic to_doctor - the specialty of the physician to whom the assistant referred the patient prompt - ready prompt for fincetuning instruct phi model… See the full description on the dataset page: https://huggingface.co/datasets/Mykes/rus_med_dialogues_qa.textquestion-answering1K<n<10K4 likes69 downloads2y agoHugging Face29gretelai /commonsense-dialogues Commonsense-Dialogues Dataset This is the Commonsense-Dialogues, a crowdsourced dataset of ~11K dialogues grounded in social contexts involving utilization of commonsense. The dataset was released by Amazon Alexa AI team in collaboration with the University of Southern California (USC), and also available Commonsense-Dialogues repo The social contexts used were sourced from the train split of the SocialIQA dataset, a multiple-choice question-answering based social commonsense… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/commonsense-dialogues.texttext-classification10K<n<100K6 likes68 downloads2y agoHugging Face30ychen /Generated-Empathetic-Dialogues-v0.1-Smol Generated Empathetic Conversations v0.1 - Smol This is dataset contains 10K rows of multi-round empathetic conversations convering a diverse set of topics. Highlights Multi-round conversation It's not single-turn. The user and the assistant works together to gradually unfold the conversation. The average number of turns is 5, with a standard deviation of approximately 1.59 turns. A turn consists of two messages with one by the user, and another by the… See the full description on the dataset page: https://huggingface.co/datasets/ychen/Generated-Empathetic-Dialogues-v0.1-Smol.texttext-generation10K<n<100K4 likes59 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.