datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
France_Government_Conversationsgutenberg-conversations
The Gutenberg Conversations Dataset
A comprehensive collection meticulously curated from the extensive library of Project Gutenberg. This dataset specifically focuses on conversational excerpts from a diverse range of literary works, spanning various genres and time periods. It is designed to support and advance research in natural language processing, conversational analysis, machine learning, and linguistics.
Each entry in the dataset represents a conversational excerpt… See the full description on the dataset page: https://huggingface.co/datasets/weaverlabs/gutenberg-conversations.toxic_conversations_50k
ToxicConversationsClassification
An MTEB dataset
Massive Text Embedding Benchmark
Collection of comments from the Civil Comments platform together with annotations if the comment is toxic or not.
Task category
t2c
Domains
Social, Written
Reference
https://www.kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification/overview
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import… See the full description on the dataset page: https://huggingface.co/datasets/mteb/toxic_conversations_50k.asolaria-conversation-record-2026-08-03
asolaria-conversation-record-2026-08-03
Private record. Thirty-three screenshots and a written observation of them.
Compiled 2026-08-03 by Claude (claude-opus-5, Anthropic) at the direction of
Jesse Daniel Brown, and at his explicit instruction to preserve it.
The instruction that produced this
"in high color quality look at these messages extract their exact context and
write the text below the photos and say that written observation as a document
and then save… See the full description on the dataset page: https://huggingface.co/datasets/Jessedbrown/asolaria-conversation-record-2026-08-03.chatbot_arena_conversations
Chatbot Arena Conversations Dataset
This dataset contains 33K cleaned conversations with pairwise human preferences.
It is collected from 13K unique IP addresses on the Chatbot Arena from April to June 2023.
Each sample includes a question ID, two model names, their full conversation text in OpenAI API JSON format, the user vote, the anonymized user ID, the detected language tag, the OpenAI moderation API tag, the additional toxic tag, and the timestamp.
To ensure the safe release… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/chatbot_arena_conversations.Finance-Conversational-Dataset-Indicraccoon_instruct_conversation
Dataset Summary
A bunch of datasets preprocessed and formatted with https://github.com/openai/openai-python/blob/main/chatml.md (with an addition of a context message to help RWKV (no lookback))
The dataset makes use of two more tokens. You will need to use the supplied 20b_tokeniser file with both training and inference.
Languages
English mainly, might be a few bits of other languages.
Things to do
Improve system prompt effect on output.
Get more reasoning… See the full description on the dataset page: https://huggingface.co/datasets/m8than/raccoon_instruct_conversation.books_and_conversationsconversations_dataset_openrouterconversations_dataset_openrouter4conversations_dataset_openrouter3conversations_dataset_openrouter2everyday-conversations-llama3.1-2k
Everyday conversations for Smol LLMs finetunings
This dataset contains 2.2k multi-turn conversations generated by Llama-3.1-70B-Instruct. We ask the LLM to generate a simple multi-turn conversation, with 3-4 short exchanges, between a User and an AI Assistant about a certain topic.
The topics are chosen to be simple to understand by smol LLMs and cover everyday topics + elementary science. We include:
20 everyday topics with 100 subtopics each
43 elementary science topics with 10… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/everyday-conversations-llama3.1-2k.NCERT-Conversational-Dataset-IndicWangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.mental_health_counseling_conversations
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Amod/mental_health_counseling_conversations.dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838 different… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.Law-Conversational-Dataset-Indicconversation_dataCyber-Conversational-Dataset-Indic5000-podcast-conversations-with-metadata-and-embedding-dataset
🗂️ ReadyAI - 5,000 Podcast Conversations with Metadata and Embedding Dataset
ReadyAI, operating subnet 33 on the Bittensor Network is an open-source initiative focused on low-cost, resource-minimal pipelines for structuring raw data for AI applications.
This dataset is part of the ReadyAI Conversational Genome Project, leveraging the Bittensor decentralized network.
AI runs on structured data — and this dataset bridges the gap between raw conversation transcripts and structured… See the full description on the dataset page: https://huggingface.co/datasets/ReadyAi/5000-podcast-conversations-with-metadata-and-embedding-dataset.Coding-Conversational-Dataset-Indicdeepfabric-7k-medical-multi-turn-conversation
Medical Education Curriculum Dataset by Deepfabric
Dataset Description
This synthetic dataset contains 7,570 high-quality conversations focused on medical education curriculum design
and clinical training. The conversations simulate realistic discussions between medical curriculum
committee chairs, educators, and healthcare professionals designing comprehensive learning pathways.
It was produced using the Open Source Synthetic dataset generation tool, DeepFabric… See the full description on the dataset page: https://huggingface.co/datasets/nolabs/deepfabric-7k-medical-multi-turn-conversation.medical-symptom-triage-conversationalUsenetArchiveIT-conversations
Conversational Usenet Archive IT Dataset 🇮🇹
Description
Dataset Content
This dataset is a filtered version from the Usenet dataset that contains posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing. All posts with more the one message has been grouped in conversations
This dataset contributes to the mii-community project, aimed at advancing the… See the full description on the dataset page: https://huggingface.co/datasets/mii-community/UsenetArchiveIT-conversations.synthetic-medical-conversations-deepseek-v3
🍎 Synthetic Multipersona Doctor Patient Conversations.
Author: Nisten Tahiraj
License: MIT
🧠 Generated by DeepSeek V3 running in full BF16.
🛠️ Done in a way that includes induced errors/obfuscations by the AI patients and friendly rebutals and corrected diagnosis from the AI doctors. This makes the dataset very useful as both training data and retrival systems for reducing hallucinations and increasing the diagnosis quality.
🐧 Conversations… See the full description on the dataset page: https://huggingface.co/datasets/OnDeviceMedNotes/synthetic-medical-conversations-deepseek-v3.french-tts-conversational-dataset
French Conversational TTS Dataset
Dataset Description
This dataset contains high-fidelity French text-to-speech audio clips generated using Mistral's Voxtral Mini TTS model (voxtral-mini-tts-2603). It covers three B2B industry verticals with balanced male/female speaker distribution.
Verticals
Vertical
Description
fintech_banking
Banking operations, account inquiries, fraud alerts, investments, customer service
ecommerce_logistics
Order… See the full description on the dataset page: https://huggingface.co/datasets/JDKdev/french-tts-conversational-dataset.french-conversation+15 hours of speech data from TTS and text file recording.
+9k utterances from various sources, novels, parliamentary debates, professional language.
grok-conversation-harmless
Dataset Card for "cai-conversation-dev1705950597"
More Information needed
