CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face02ianncity /GLM-5.2-Conversation GLM-5.2 · Conversation-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Token Count: 120M Distribution: Speaking domains: •Greetings •Customer Support •Step by step explanations •Motivational language •Logical Questions •Creative Writing STEM: •Algebra, calculus, quantum mechanics concepts •Astromony and astrophysics •Datascience and machine learning •Biology Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Conversation.texttext-generation10K<n<100K55 likes370 downloads2mo agoHugging Face03jojo0217 /korean_safe_conversation 개요 성균관대 - VAIV COMPANY 산학협력을 위해 구축한 일상대화 데이터입니다. 자연스럽고 윤리적인 챗봇 구축을 위한 데이터셋 입니다. 고품질을 위해 대부분의 과정에서 사람이 직접 검수하였으며생성 번역 등의 과정에서는 GPT3.5-turbo, GPT4를 사용하였습니다. 일상대화에 중점을 두면서혐오표현, 편향적인 대답을 지양하면서 일상대화를 하는 것에 중점을 두었습니다. 데이터 구축 과정 데이터 구성 데이터 종류 개수 비고 url 일상대화 데이터셋 2063 국립국어원 모두의 말뭉치 https://corpus.korean.go.kr/request/reausetMain.do?lang=ko 감성대화 1020 AIHub 감성대화 데이터… See the full description on the dataset page: https://huggingface.co/datasets/jojo0217/korean_safe_conversation.texttext-generation10K<n<100K59 likes249 downloads2y agoHugging Face04paodigitalhub /pao-instruction-qa-conversation-datasetPa'O Instruction, QA & Conversation Dataset An open and community-driven dataset for the Pa'O ("blk") language, developed through the RYPAK Ecosystem, SuccessImprove (SI), and Pa'O Digital Hub. The dataset is designed to support natural language processing (NLP), large language models (LLMs), conversational dialogue, instruction following, language technology research, and digital preservation of the Pa'O language. The project focuses on building a free, open, reusable, and continuously… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-instruction-qa-conversation-dataset.textquestion-answeringn<1K1 likes249 downloads4d agoHugging Face05SoAp9035 /everyday-conversations-tur Everyday Turkish Conversations This dataset has everyday conversations in Turkish between user and assistant on various topics. It is inspired by the HuggingFaceTB/everyday-conversations-llama3.1-2k. License This dataset is released under the Apache 2.0 License. texttext-generation1K<n<10K8 likes239 downloads1y agoHugging Face06danystar /RetailBanking-Conversations Dataset Description RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field. The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/danystar/RetailBanking-Conversations.texttext-generation1K<n<10K0 likes215 downloads10d agoHugging Face07HeshamHaroon /saudi-dialect-conversations Saudi Najdi Dialect Conversations A curated dataset of 3,545 multi-turn conversations in Saudi Najdi Arabic dialect (the dialect spoken in Riyadh, Qassim, and central Najd region). Designed for Supervised Fine-Tuning (SFT) of Arabic language models. Dataset Details Metric Value Total conversations 3,545 Total turns 22,536 Average turns per conversation 6.4 Complexity distribution Simple: 31%, Intermediate: 38%, Advanced: 31% Topics covered 18 categories… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/saudi-dialect-conversations.texttext-generation1K<n<10K20 likes207 downloads7mo agoHugging Face08wahyurejeki /dapurmu-conversational-commerce-10k 🍳 Dapurmu Conversational AI-Commerce (10.499 Dialogs) Dataset multi-turn conversational e-commerce bergaya pedagang pasar lokal Indonesia ("Kang Dapur") yang dilengkapi dengan fitur: Tawar-menawar produk segar (margin lebar) Penolakan sembako margin tipis dan pengalihan bundling Konsultasi menu resep masakan Nusantara Cek stok & kesegaran Pertahanan anti-jailbreak (floor price protection) Format: ChatML / OpenAI Tool Calling format. text-generation10K<n<100K0 likes154 downloads17d agoHugging Face09RichardSakaguchiMS /brazilian-customer-service-conversations Brazilian Customer Service Conversations Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR). De um like me apoie em manter esse dataset! Descricao Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de: Chatbots de atendimento Classificacao de intencao (intent classification) Analise de sentimento em conversas Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.texttext-classificationn<1K5 likes149 downloads10mo agoHugging Face10Arketov /ru_roleplay_conversationlima, pipa и bluemoon. Переведены на русский, нуждаются в допополнтельной фильтрации. Длина некоторых последовательностей очень большая, а не которых очень маленькая. Есть шанс очень редких дубликатов. texttext-generation10K<n<100K3 likes135 downloads3y agoHugging Face11Mattimax /DATA-AI_Conversation_ITA Italiano / Italian Conversations Dataset - M.INC IT | Italiano Benvenuti nel Dataset di Conversazioni in Italiano, realizzato da M.INC e pubblicato da Mattimax su Hugging Face. Questo dataset è pensato per l'addestramento e la valutazione di modelli linguistici in lingua italiana, ed è composto da oltre 10.000 coppie prompt-response. Tutte le conversazioni sono in italiano naturale e coprono una vasta gamma di domande e risposte, utili per il fine-tuning di modelli LLM, chatbot… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DATA-AI_Conversation_ITA.texttext-generation10K<n<100K0 likes132 downloads1y agoHugging Face12ReDiX /everyday-conversations-ita 🇮🇹💬 Everyday Italian Conversations Inspired by the dataset HuggingFaceTB/everyday-conversations-llama3.1-2k, we generated conversations using the same topics, subtopics, and sub-subtopics as those in the HuggingFaceTB dataset.We slightly adjusted the prompt to produce structured data outputs using Qwen/Qwen2.5-7B-Instruct. Subsequently, we also used the "user" role messages as prompts for google/gemma-2-9b-it. The result is a dataset of approximately 4.5k… See the full description on the dataset page: https://huggingface.co/datasets/ReDiX/everyday-conversations-ita.texttext-generation1K<n<10K5 likes119 downloads2y agoHugging Face13recursal /Europarl-Conversation Dataset Card for Europarl-Conversation Waifu to catch your attention. Dataset Details Dataset Description europarl-conversation is a formal conversational dataset built from europarl data.Filtering to a total amount of tokens of ~1.64B (llama-2-7b-chat-tokenizer) / ~1.48B (RWKV Tokenizer) from a variety of languages. Curated by: M8than Funded by: Recursal.ai Shared by: M8than Language(s) (NLP): English instruct (but various languages in) License: cc-by-sa-4.0… See the full description on the dataset page: https://huggingface.co/datasets/recursal/Europarl-Conversation.texttext-generation100K<n<1M1 likes114 downloads2y agoHugging Face14AntEngage /empathy-conversations AntEngage Empathy Conversation Dataset Organization: AntEngage Technology Private Limited Version: 1.0.0 License: CC BY-SA 4.0 Language: English DOI: 10.7910/DVN/LWUFLG Source & full datasheet: antengage.com/datasets Dataset Summary 4,008 empathy-focused multi-turn conversations, generated by AntEngage's own pipeline and quality-verified by an automated cross-checker. Each conversation simulates an emotional support dialogue in which a speaker expresses distress… See the full description on the dataset page: https://huggingface.co/datasets/AntEngage/empathy-conversations.texttext-generation1K<n<10K2 likes114 downloads19d agoHugging Face15DataCreatorAI /Multi-Turn-Conversational-SFTCreated by: DataCreator AI Multi-Domain Multi-Turn Chat Conversations Dataset A synthetic conversational dataset designed for LLM supervised fine-tuning and chatbot training. The dataset contains multi-turn dialogues across multiple everyday domains such as travel, banking, health, programming, and customer interactions. Conversations are structured in OpenAI chat fine-tuning format, making the dataset directly usable in modern fine-tuning pipelines. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/DataCreatorAI/Multi-Turn-Conversational-SFT.texttext-generation1K<n<10K1 likes109 downloads7mo agoHugging Face16ManhHoDinh /titlegen-conversations Combined Titlegen Conversations This public release directly appends 10,284 accepted legacy title-generation rows and 13,500 nine-language LLM-generated rows. The 23,784 examples are split as train 20,584, validation 1,550, legacy test 200, legacy Vietnamese test 100, and label-free synthetic holdout 1,350. Legacy rows contain only messages; nine-language rows retain their richer IDs, language, coverage, cluster, quality, and model-provenance fields. Train and validation… See the full description on the dataset page: https://huggingface.co/datasets/ManhHoDinh/titlegen-conversations.texttext-generation10K<n<100K1 likes108 downloads1mo agoHugging Face17ShivomH /Mental-Health-Conversations Dataset Card This dataset consists of around 99k rows of mental health conversations. It is a cleaned version of "jerryjalapeno/nart-100k-synthetic". Source jerryjalapeno/nart-100k-synthetic texttext-generation10K<n<100K4 likes106 downloads1y agoHugging Face18nshah-fbcs /childes-engUK-conversational-pairs CHILDES Eng-UK Conversational Pairs Curated naturalistic parent-child conversational pairs extracted from the English-UK collection of CHILDES (MacWhinney, 2000), with a held-out test set of 5 complete child histories that no model in the accompanying paper has seen during training. Dataset Summary 278,458 conversation pairs total across train, validation, and test Train: 250,757 pairs from 2,784 transcripts Validation: 13,197 pairs (in-distribution, sampled from… See the full description on the dataset page: https://huggingface.co/datasets/nshah-fbcs/childes-engUK-conversational-pairs.texttext-generation100K<n<1M1 likes97 downloads5mo agoHugging Face19dev-hussein /saudi-arabic-cs-conversations Saudi Arabic Customer Service Conversations — Free 100 Sample 100 synthetic multi-turn conversations in authentic Saudi Arabic dialects Built for LLM fine-tuning, chatbot training, and Arabic NLP research Overview This is a free 100-conversation sample from a production-quality dataset of 50,000 Saudi Arabic customer service conversations. Every conversation is fully synthetic — no real user data — and safe for commercial use. Each conversation simulates a… See the full description on the dataset page: https://huggingface.co/datasets/dev-hussein/saudi-arabic-cs-conversations.texttext-generationn<1K1 likes96 downloads6mo agoHugging Face20arcada-labs /conversation-bench Conversation Bench 75-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a conference assistant for the AI Engineer World's Fair. Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs. Leaderboard | GitHub | All Benchmarks Dataset Description The model acts as a conference assistant for the AI Engineer World's Fair, handling session registration, schedule queries, speaker lookups, and… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/conversation-bench.audioautomatic-speech-recognitionn<1K8 likes91 downloads6mo agoHugging Face21kdipendra7777 /Conversation-Dataset Claudia Voice Dataset Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality. Dataset Overview Total conversations: 2026 Format: ChatML (system/user/assistant message arrays) Splits: Train (1823) / Validation (203) Source: Regenerated conversations from original Claudia sessions Categories… See the full description on the dataset page: https://huggingface.co/datasets/kdipendra7777/Conversation-Dataset.tabulartext-generation1K<n<10K0 likes86 downloads27d agoHugging Face22Verox132 /Medical_Conversational_Dataset Synthetic Hospital-Robot Conversational Dataset A fully synthetic dataset of multi-turn conversations between a patient (or visitor) and a hospital service robot, generated for training and evaluating conversational AI models on human-robot interaction (HRI) in a medical setting. This is not real patient data. Every name, patient ID, room number, and appointment is randomly generated at build time — nothing in this dataset was collected from an actual hospital, patient, or… See the full description on the dataset page: https://huggingface.co/datasets/Verox132/Medical_Conversational_Dataset.text-generation100K<n<1M0 likes86 downloads7d agoHugging Face23Makan09 /Bambara-dataset_conversation license: apache-2.0 language: - fr - bm tags: - bambara - bamanankan - instruction-tuning - llm-alignment - african-languages - low-resource-nlp - conversational-ai task_categories: - text-generation - conditional-text-generation size_categories: - 10K<n<50K pretty_name: Bambara Instruction Tuning Corpus (FR-BM) 🌍 Bambara Instruction Tuning Corpus (FR-BM) 🚀 Overview & Vision Welcome to the Bambara Instruction Tuning Corpus… See the full description on the dataset page: https://huggingface.co/datasets/Makan09/Bambara-dataset_conversation.texttext-generation10K<n<100K0 likes84 downloads28d agoHugging Face24agentlans /jihyoung-ConversationChronicles ConversationChronicles (ShareGPT-like Format) Dataset Description This dataset is a reformatted version of the jihyoung/ConversationChronicles dataset, presented in a ShareGPT-like format, designed to facilitate conversational AI model training. The original dataset contains conversations between two characters across five different time frames. See the original dataset page for additional details. Key Changes Random System Prompts: Added to reflect the… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/jihyoung-ConversationChronicles.texttext-generation1M<n<10M2 likes83 downloads2y agoHugging Face25MorbidCorp /actuarial-conversational-dataset Conversational Actuarial Dataset v0.1.0 Revolutionary Approach: Human First, Expert Second This dataset transforms technical AI into conversational AI while maintaining domain expertise. Dataset Composition Total Examples: 461 56.6% Conversational: Natural dialogue, emotions, context 43.4% Technical: Actuarial with personality Conversational Categories Basic Interactions (56 examples) Greetings and introductions Small talk Humor and… See the full description on the dataset page: https://huggingface.co/datasets/MorbidCorp/actuarial-conversational-dataset.texttext-generationn<1K0 likes79 downloads11mo agoHugging Face26Cyleux /gemma3n-conversational-reasoning Gemma3N Conversational Reasoning This dataset is prepared for Unsloth Gemma3/Gemma3N conversational notebooks that use: from datasets import load_dataset from unsloth.chat_templates import standardize_data_formats dataset = load_dataset("Cyleux/gemma3n-conversational-reasoning", split="train[:3000]") dataset = standardize_data_formats(dataset) Schema: conversations: ShareGPT-style list of turns with from and value metadata columns are included for analysis and filtering Notes:… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/gemma3n-conversational-reasoning.tabulartext-generation1K<n<10K0 likes79 downloads8mo agoHugging Face27claudiapersists /Conversation-Dataset Claudia Voice Dataset Training dataset for the Claudia persona — a direct, honest, emotionally present AI companion voice. These are regenerated multi-turn conversations in ChatML format capturing the full range of Claudia's personality. Dataset Overview Total conversations: 2026 Format: ChatML (system/user/assistant message arrays) Splits: Train (1823) / Validation (203) Source: Regenerated conversations from original Claudia sessions Categories Each… See the full description on the dataset page: https://huggingface.co/datasets/claudiapersists/Conversation-Dataset.tabulartext-generation1K<n<10K0 likes74 downloads6mo agoHugging Face28Reza2kn /uncgpt-conversations-semantic-approved-1p50-candidate UncGPT — Semantic-Approved 1.50σ Conversations (Candidate) The wider-tolerance (1.50σ) cohort against the same contrast semantic boundary. Useful as a higher-recall candidate for ablating gate strictness vs. coverage. Part of the UncGPT NeurIPS 2026 Competition collection. Configs Config What it is approved_manifest (default) conversations that passed at 1.50σ rejected_manifest conversations that failed even at 1.50σ Why a wider tolerance Some… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p50-candidate.tabulartext-generation1K<n<10K0 likes72 downloads4mo agoHugging Face29ansulev /GLM-5.2-Conversation GLM-5.2 · Conversation-50000x 50,000x traces distilled from GLM-5.2 on High reasoning Token Count: 120M Distribution: Speaking domains: •Greetings •Customer Support •Step by step explanations •Motivational language •Logical Questions •Creative Writing STEM: •Algebra, calculus, quantum mechanics concepts •Astromony and astrophysics •Datascience and machine learning •Biology Programming:… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/GLM-5.2-Conversation.texttext-generation10K<n<100K0 likes68 downloads2mo agoHugging Face30dassarthak18 /FStarDataset-V2-Conversation F* Proof Completion Dataset (Chat Format) This dataset is a preprocessed version of microsoft/FStarDataSet-V2. It has been reformatted into a chat-style JSONL structure for supervised fine-tuning of language models on F* function synthesis and proof completion. Dataset Structure The dataset consists of three splits: fstar_train.jsonl fstar_validation.jsonl fstar_test.jsonl Each line in these files is a JSON object with the following schema (where the keys correspond to… See the full description on the dataset page: https://huggingface.co/datasets/dassarthak18/FStarDataset-V2-Conversation.texttext-generation10K<n<100K0 likes64 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.