CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01weaverlabs /gutenberg-conversations The Gutenberg Conversations Dataset A comprehensive collection meticulously curated from the extensive library of Project Gutenberg. This dataset specifically focuses on conversational excerpts from a diverse range of literary works, spanning various genres and time periods. It is designed to support and advance research in natural language processing, conversational analysis, machine learning, and linguistics. Each entry in the dataset represents a conversational excerpt… See the full description on the dataset page: https://huggingface.co/datasets/weaverlabs/gutenberg-conversations.text10K<n<100K1 likes3.4k downloads2y agoHugging Face02mteb /toxic_conversations_50k ToxicConversationsClassification An MTEB dataset Massive Text Embedding Benchmark Collection of comments from the Civil Comments platform together with annotations if the comment is toxic or not. Task category t2c Domains Social, Written Reference https://www.kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification/overview How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import… See the full description on the dataset page: https://huggingface.co/datasets/mteb/toxic_conversations_50k.texttext-classification100K<n<1M19 likes3.1k downloads7mo agoHugging Face03lmsys /chatbot_arena_conversationsgated Chatbot Arena Conversations Dataset This dataset contains 33K cleaned conversations with pairwise human preferences. It is collected from 13K unique IP addresses on the Chatbot Arena from April to June 2023. Each sample includes a question ID, two model names, their full conversation text in OpenAI API JSON format, the user vote, the anonymized user ID, the detected language tag, the OpenAI moderation API tag, the additional toxic tag, and the timestamp. To ensure the safe release… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/chatbot_arena_conversations.tabular10K<n<100K490 likes2.4k downloads3y agoHugging Face04oss-codes /Finance-Conversational-Dataset-Indictext100K<n<1M1 likes2.3k downloads1y agoHugging Face05giannisan /conversations_dataset_openrouter4textn<1K0 likes2.1k downloads2y agoHugging Face06giannisan /conversations_dataset_openroutertextn<1K0 likes2.1k downloads2y agoHugging Face07giannisan /conversations_dataset_openrouter3textn<1K0 likes2.1k downloads2y agoHugging Face08giannisan /conversations_dataset_openrouter2textn<1K0 likes2.1k downloads2y agoHugging Face09HuggingFaceTB /everyday-conversations-llama3.1-2k Everyday conversations for Smol LLMs finetunings This dataset contains 2.2k multi-turn conversations generated by Llama-3.1-70B-Instruct. We ask the LLM to generate a simple multi-turn conversation, with 3-4 short exchanges, between a User and an AI Assistant about a certain topic. The topics are chosen to be simple to understand by smol LLMs and cover everyday topics + elementary science. We include: 20 everyday topics with 100 subtopics each 43 elementary science topics with 10… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k.text1K<n<10K138 likes2k downloads2y agoHugging Face10ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face11Amod /mental_health_counseling_conversationsgated Amod/mental_health_counseling_conversations This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue. Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Amod/mental_health_counseling_conversations.texttext-generation1K<n<10K502 likes1.7k downloads10mo agoHugging Face12langswap /dialogs-ru-emotional-conversations Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational Russian speech, designed for dialog-oriented and emotional text-to-speech. Unlike existing Russian corpora — mostly single-speaker read speech or large but low-quality web-mined audio — Dialogs was recorded by professional theatre actors performing scripted dialogs face-to-face, capturing natural turn-taking, timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.audiotext-to-speechn<1K18 likes1.6k downloads2mo agoHugging Face13oss-codes /Law-Conversational-Dataset-Indictext100K<n<1M0 likes1.4k downloads1y agoHugging Face14nvidia /Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 Dataset Description: We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838 different… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.tabular10K<n<100K32 likes1.4k downloads7mo agoHugging Face15oss-codes /Cyber-Conversational-Dataset-Indictext1K<n<10K0 likes1.3k downloads1y agoHugging Face16oss-codes /NCERT-Conversational-Dataset-Indictext100K<n<1M0 likes1.3k downloads1y agoHugging Face17ReadyAi /5000-podcast-conversations-with-metadata-and-embedding-dataset 🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications. This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network. AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.text10K<n<100K8 likes1.2k downloads1y agoHugging Face18nolabs /deepfabric-7k-medical-multi-turn-conversation Medical Education Curriculum Dataset by Deepfabric Dataset Description This synthetic dataset contains 7,570 high-quality conversations focused on medical education curriculum design and clinical training. The conversations simulate realistic discussions between medical curriculum committee chairs, educators, and healthcare professionals designing comprehensive learning pathways. It was produced using the Open Source Synthetic dataset generation tool, DeepFabric… See the full description on the dataset page: https://huggingface.co/datasets/nolabs/deepfabric-7k-medical-multi-turn-conversation.text1K<n<10K1 likes1.1k downloads11mo agoHugging Face19sweatSmile /medical-symptom-triage-conversationaltabular1K<n<10K1 likes1.1k downloads1y agoHugging Face20OnDeviceMedNotes /synthetic-medical-conversations-deepseek-v3 🍎 Synthetic Multipersona Doctor Patient Conversations. Author: Nisten Tahiraj License: MIT 🧠 Generated by DeepSeek V3 running in full BF16. 🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival systems for reducing hallucinations and increasing the diagnosis quality. 🐧 Conversations… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/synthetic-medical-conversations-deepseek-v3.text35 likes1k downloads2y agoHugging Face21mii-community /UsenetArchiveIT-conversations Conversational Usenet Archive IT Dataset 🇮🇹 Description Dataset Content This dataset is a filtered version from the Usenet dataset that contains posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing. All posts with more the one message has been grouped in conversations This dataset contributes to the mii-community project, aimed at advancing the… See the full description on the dataset page: https://huggingface.co/datasets/mii-community/UsenetArchiveIT-conversations.texttext-generation1M<n<10M8 likes973 downloads2y agoHugging Face22MaziyarPanahi /synthetic-medical-conversations-deepseek-v3-chatTaken from Synthetic Multipersona Doctor Patient Conversations. by Nisten Tahiraj. Original README 🍎 Synthetic Multipersona Doctor Patient Conversations. Author: Nisten Tahiraj License: MIT 🧠 Generated by DeepSeek V3 running in full BF16. 🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/synthetic-medical-conversations-deepseek-v3-chat.text1K<n<10K6 likes880 downloads2y agoHugging Face23ElementXMaster /conversational_raw_datasettext10M<n<100M0 likes803 downloads1y agoHugging Face24oss-codes /Computer-Science-Conversational-Dataset-Indictext10K<n<100K0 likes784 downloads1y agoHugging Face25AdaMLLab /smolkalam-arabic-conversational-sft SmolKalam SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets. Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.tabulartext-generation1M<n<10M3 likes770 downloads1mo agoHugging Face26TNE-AI /customer-support-on-twitter-conversationtext100K<n<1M6 likes753 downloads2y agoHugging Face27DeepMostInnovations /saas-sales-conversations saas-sales-conversations Dataset Description This is a synthetic dataset of sales conversations for SaaS (Software as a Service) companies, designed for training sales conversion prediction models. The dataset was created following the methodology presented in "SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization" (Nandakishor M, 2025). The dataset contains realistic dialogues between sales representatives and… See the full description on the dataset page: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations.tabulartext-classification100K<n<1M46 likes728 downloads1y agoHugging Face28health360 /Ultrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized Dataset Card for "Ultrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized" More Information needed text1M<n<10M0 likes658 downloads3y agoHugging Face29oss-codes /CA-Conversational-Dataset-Indictext100K<n<1M0 likes626 downloads1y agoHugging Face30B-Lounes /conversational_audio_fr_dataset-metadatatextn<1K0 likes605 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.