datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chatbot_arena_conversations
Chatbot Arena Conversations Dataset
This dataset contains 33K cleaned conversations with pairwise human preferences.
It is collected from 13K unique IP addresses on the Chatbot Arena from April to June 2023.
Each sample includes a question ID, two model names, their full conversation text in OpenAI API JSON format, the user vote, the anonymized user ID, the detected language tag, the OpenAI moderation API tag, the additional toxic tag, and the timestamp.
To ensure the safe release… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/chatbot_arena_conversations.dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.conversations_dataset_openrouter2conversations_dataset_openrouterconversations_dataset_openrouter4conversations_dataset_openrouter3everyday-conversations-llama3.1-2k
Everyday conversations for Smol LLMs finetunings
This dataset contains 2.2k multi-turn conversations generated by Llama-3.1-70B-Instruct. We ask the LLM to generate a simple multi-turn conversation, with 3-4 short exchanges, between a User and an AI Assistant about a certain topic.
The topics are chosen to be simple to understand by smol LLMs and cover everyday topics + elementary science. We include:
20 everyday topics with 100 subtopics each
43 elementary science topics with 10… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k.5000-podcast-conversations-with-metadata-and-embedding-dataset
🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset
ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications.
This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network.
AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.deepfabric-7k-medical-multi-turn-conversation
Medical Education Curriculum Dataset by Deepfabric
Dataset Description
This synthetic dataset contains 7,570 high-quality conversations focused on medical education curriculum design
and clinical training. The conversations simulate realistic discussions between medical curriculum
committee chairs, educators, and healthcare professionals designing comprehensive learning pathways.
It was produced using the Open Source Synthetic dataset generation tool, DeepFabric… See the full description on the dataset page: https://huggingface.co/datasets/nolabs/deepfabric-7k-medical-multi-turn-conversation.medical-symptom-triage-conversationalUsenetArchiveIT-conversations
Conversational Usenet Archive IT Dataset 🇮🇹
Description
Dataset Content
This dataset is a filtered version from the Usenet dataset that contains posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing. All posts with more the one message has been grouped in conversations
This dataset contributes to the mii-community project, aimed at advancing the… See the full description on the dataset page: https://huggingface.co/datasets/mii-community/UsenetArchiveIT-conversations.synthetic-medical-conversations-deepseek-v3-chatTaken from Synthetic Multipersona Doctor Patient Conversations. by Nisten Tahiraj.
Original README
🍎 Synthetic Multipersona Doctor Patient Conversations.
Author: Nisten Tahiraj
License: MIT
🧠 Generated by DeepSeek V3 running in full BF16.
🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival… See the full description on the dataset page: https://huggingface.co/datasets/MaziyarPanahi/synthetic-medical-conversations-deepseek-v3-chat.smolkalam-arabic-conversational-sft
SmolKalam
SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets.
Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.conversational_raw_datasetcustomer-support-on-twitter-conversationUltrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized
Dataset Card for "Ultrachat-Multiple-Conversations-Alpaca-Tinyllama-Tokenized"
More Information needed
lmsys-chatbot_arena_conversations
Dataset Card for "lmsys-chatbot_arena_conversations"
More Information needed
expresso-conversational
The Expresso Dataset
[paper] [demo samples] [Original repository]
Introduction
The Expresso dataset is a high-quality (48kHz) expressive speech dataset that includes both expressively rendered read speech (8 styles, in mono wav format) and improvised dialogues (26 styles, in stereo wav format). The dataset includes 4 speakers (2 males, 2 females), and totals 40 hours (11h read, 30h improvised). The transcriptions of the read speech are also provided.
You can listen to… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/expresso-conversational.FinePersonas-Synthetic-Email-Conversations
FinePersonas Synthetic Email Conversations
FinePersonas Synthetic Email Conversations is a dataset containing around 115k conversations via email between two personas from the argilla/FinePersonas-v0.1. Conversations were generated using NousResearch/Hermes-3-Llama-3.1-70B.
🗞️ News
[10/16/2024] New subsets: added two new subsets unfriendly_email_conversations and unprofessional_email_conversations.
How were the conversations generated?… See the full description on the dataset page: https://huggingface.co/datasets/argilla/FinePersonas-Synthetic-Email-Conversations.conversational_data_untokenized_mergedcai-conversation-harmless
Dataset Card for "cai-conversation-dev1705629166"
More Information needed
Ultrachat-Multiple-Conversations-Alpaca-Style
Dataset Card for "Ultrachat-Multiple-Conversations-Alpaca-Style"
More Information needed
Domofon-Cot-Conversations-700k
Domofon-Cot-Conversations-700k
Synthetic XML conversation data for training small language models on reasoning,
instruction following, XML formatting, and tool-use traces.
Repository: domofon/Domofon-Cot-Conversations-700k
What is inside
The dataset contains cleaned generated XML conversations from six families:
conv: multi-turn factual conversations with tool-use traces.
instruct: text-processing instructions, including deterministic count tool calls.
ds:… See the full description on the dataset page: https://huggingface.co/datasets/domofon/Domofon-Cot-Conversations-700k.SynSQL-2.5M-conversations-standardizedhuman_assistant_conversationsales-conversations
Dataset Card for "sales-conversations"
This dataset was created for the purpose of training a sales agent chatbot that can convince people.
The initial idea came from: textbooks is all you need https://arxiv.org/abs/2306.11644
gpt-3.5-turbo was used for the generation
Structure
The conversations have a customer and a salesman which appear always in changing order. customer, salesman, customer, salesman, etc.
The customer always starts the conversation
Who ends the… See the full description on the dataset page: https://huggingface.co/datasets/goendalf666/sales-conversations.en-uk-translation-conversationsdeepa2-conversations
Summary
This dataset contains multi-turn conversations that gradually unfold deep logical analyses of argumentative texts.
In particular, the chats contain examples of how to
use Argdown syntax
logically formalize arguments in FOL (latex, nltk etc.)
annotate an argumentative text
use Z3 theorem prover to check deductive validity
use custom tools in conjunction with argument reconstructions
The chats are template-based renderings of the synthetic, comprehensive argument analyses… See the full description on the dataset page: https://huggingface.co/datasets/DebateLabKIT/deepa2-conversations.hle-no-img-conversational-formatmental_health_counseling_conversations_sharegpt
Dataset Card for "mental_health_counseling_conversations_sharegpt"
More Information needed
